Sad reality, I'm using LLMs to review the work of our offshore programmers, and as you say: some are good, some are not - and sadly, the ones who are not are also not learning to leverage LLMs to improve their work products. Without LLM review, I'd be advocating to find some of our offshore programmers "more productive ways to apply their skills." With four rounds of LLM review, I'm effectively rewriting (the bad ones') code for them... every... single... pull request.
My best offshore coder figured out how to use the LLMs to review his own code, we have architecture discussions about the best things to do and never have issues with how they are done.
My worst offshore coder "doesn't believe in using AI" and continues to submit code for merge to master branch with obvious race conditions, panic crashes, etc.
A lot has been made about the tokens being subsidized, and at the frontier models I believe that's very true. I also believe that we're getting to where the frontier may be better, but not necessary, in order to get value out of using the LLM - and a couple of steps back from the frontier is becoming quite affordable now.
I'd be very interested to try out a GLM-5.2 instance on about a $100K server (8 bit quantitized) - which should cost under $20K per year to operate and maintain (although with component prices inflating the way they are that number continues to climb)... if that setup is as useful as, say, Claude Opus 4.5, that's a reasonable price level for a system that should be able to serve a department of 6-10 users pretty well. If GLM-5.2 isn't "there yet" - it's likely just a matter of time before one gets there.