Re: Came here to see ...
If only all that revolutionary always right tech was available when people are willing to pay lots of money, because my employers pay the big AI companies for big AI models, and we don't get that. For example, it winning top coding competitions. For one thing, there's not a lot of actual coding competitions. There are several metacoding competitions like the obfuscated C contest, code golf, etc. There are some hackathons that are open to the public. But the important thing is that people there are developing different things. There aren't many contests that actually test one programmer against another, and the main reason is that few people with skills would compete, because people hiring programmers want people who can either do a good enough job or can do something in a particularly tricky area which requires a lot more specialized knowledge. In real life, we have code generation, and we have cleaning up after it. If it was so good, why do basic employees have to read over and correct it when it's writing small utilities? Some of its output is valid. That's far different from the quality you claim.
Handling language. That's great. And so many too. So it should do a pretty good job at translation, right. I mean probably not literary translation; that's tricky, but translating some simple factual statements should work. As it happens, I also got a chance to see that in action recently, because I was localizing something to French, which I don't speak. The person who was going to do the translations was delayed, so I made the first version with AI translation as a stopgap. What did she say when she reviewed that? "This is useless, I've done it from scratch." Before you suggest it, this was not her trying to keep her job, because this was an open source project for which neither she nor I was paid a thing. And French is a language with plenty of training data. Language translation is fine for understanding a website you want to read, but if it's not good enough for translation of simple sentences in a common language, why should I expect it to do well with one with little training data which nobody at the AI company is qualified to judge?
And on that Math Olympiad performance, if that problem solving ability is so strong, why can't we run that model? It hasn't been released. I'm not actually sure what I can do with that anyway, but if I come up with a use case, I can't run the model that's capable of it. This is an issue because last year, similarly confident statements showed up claiming that a silver medal performance had been won at last year's Olympiad. What actually happened? The silver medal was truly and fairly won as long as the model didn't have to comply with the time limit and got some help parsing from some professional adult humans working in AI who understand both complex mathematics and how to prompt their LLM well. The articles I've seen suggest that the time limit was in play this year, but they're not too clear on what other conditions the thing had, and since you can speed up the model by throwing more computing at it, I have reason to ask. GPT5, on the other hand, isn't generating valid proofs when I ask for them. If I find a use for a proof-generating machine, I don't have one, and I'm wondering if maybe OpenAI doesn't really have a good one either.