Some models do use open datasets and largely reproducible training, with the Nvidia Nemotron series being the highest profile example. You are right that many are opaque about training data, though, and this is not okay.
Using them is just as unethical as using the models from OpenAI.
And I vehemently disagree with this.
OpenAI is a different order of magnitude of moral depravity. Thats like saying using Lemmy is just as bad as using Facebook or Palantir.
I dont think any spitting is necessary, nor is it productive.