AI models keep getting better

This September was a particularly busy month for AI. Fable 5.1 already feels like a distant memory, with GPT-6 Astra, Opus 5.5, Sonnet 5.5, and then GPT-6.1 Sol released shortly after. And to cap it off, Google announced Gemini 4 Argon, which at least by its shared benchmarks, suggest a return to the frontier.

To me as a consumer of AI models, this has been a (mostly) joyous and productive period. AI has been genuinely useful for me at work, especially in the realm of data analysis, as I could self-service a lot of analytics. Questions that would have been shelved a couple of years ago – due to bandwidth prioritization of scarce data analysts – now have a directional answer in 30 minutes. Having different models / harnesses independently tackle the question, and/or do adversarial reviews of others’ work, provide some reassurance that I’m not being blindly led on.

I’ve also been amused (and somewhat excited) by the discovered new capabilities of these new models. First it was GPT-6 Astra and Blender. And then Opus 5.5 showed its knack at pixel art and programmatic video generation. And of course, there’s been a flood of one-shot “make a video game” examples, which looks better and better with each new model.

For video games, there is a big difference between “looks good” and “plays good”. To quote Simon Willison (whose writings on AI I’ve greatly enjoyed):

Something I’ve realized about game development is that you can vibe-code something that looks like a computer game, and that’s easy.
Building a game that’s fun, has a good gameplay loop, and is challenging and interesting and keeps people coming back for more… that’s still beyond me, and beyond any of the agents I’ve tried.
This ties into the Deep Blue thing. Just because we can make something that looks like a game does not mean that we are game developers.

(If you are curious what Willison means by “Deep Blue”, you should read the rest of his post.)

I’d like to push this thread one step further though – what is stopping AI models from being great game developers right now? Let’s think through this.

Is it a lack of domain knowledge and theory? Likely not. The frontier models are a compression of the internet (not to mention being literal devourers of large volumes of books). On any game dev discipline, they are probably already better equipped in knowledge than the average person (or even professional game developers not of that discipline – i.e. AI is a better game artist than a typical human game programmer).

Is it due to AI outputs being too derivative? I’d give this a qualified yes. Without proper guidance, AI outputs do tend to look too familiar or over-index on established tropes (based on its training data), but even this longstanding point of criticism is weakening when the base quality of the output keeps getting better.

How about limitations in tool-use? This is where I’d argue that we’ve seen the biggest leaps in model capabilities recently, but also where it’s still lacking the most. AI is already better at web search than humans, and they can decently use browsers for various automation. More recently, we’ve been wowed by their abilities to utilize Blender, three.js, or vanilla JS (for pixel art), as well as python + ffmpeg for video. But if you’ve observed their workflow in these tasks, you’ll notice still that we’ve barely made any improvement on how AI review their work – by and large, they are still limited to taking screenshots, and then inspecting these screenshots.

Processing visual info frame-by-frame is computationally expensive, and this is easily one of the biggest areas where your token budget on frontier models get devoured. It also means that models are currently woefully ill-equipped at iterating on moment-by-moment game feel. Remember this is not just a visual task – a combat animation can look right but plays wrong, so it is the orchestration of the physical button presses and the responding audio-visual frames.

Some of this can be mitigated by making the task easier to tackle for AI. One specific path is being data-driven. If we could structure the data, such that models can see perfectly through data a comprehensive representation of the game state, then it would be akin to models having access to the firehose in the Matrix. For some tasks, this is already the established best practice – for example, having AI run lots of headless simulations to rapidly and exhaustively verify the effect of game balance changes is a powerful tool that many teams already employ.

But even if we had data abstraction such that models obtain “realtime vision” (however that’s achieved), models would still need to internalize taste to be able to evaluate and iterate towards the desired game feel. And thus we finally mention taste – perhaps the most cliched word for me in AI discussion this year. I now loathe to say it, but at least for now I can’t work around it. As long as humans are the consumers (scary thought if we aren’t), human taste is part of the equation to solve for commercial value. Any commercially successful game by definition has a commercially validated game vision. Such game visions don’t come from thin air; they are the culmination of the developers’ past experiences, judgment and taste. And materializing a game vision from a design doc requires lots of further judgment and applications of taste. In our AI-native world, we can feed AI models our design docs, and (through lots of iteration) help it refine its own heuristics to know “what good looks like for this specific team and game vision”. But it’s hard to imagine a workflow where AI models independently arrive at that end point without lots of human intervention.

So, I guess where I stand is this: can AI models be great game developers independently? No. Can they be great game developers working alongside human developers? Absolutely, and they already are.

I should stop here – this post has gone down a branch longer than I anticipated. But back to the main trunk of observing AI models getting better rapidly – there is one more thing to mention. I believe we are seeing real world examples of the “lightspeed leapfrog” trope from sci-fi. AI models are gaining specific capabilities faster than people building harnesses for specific use cases. The best-in-class pixel art workflow 6 months ago might involve a mix of 2d image generation and feeding images into video generators (seedance etc.); with Opus 5.5 and beyond, maybe the workflow has migrated to programmatic generation. So it’s quite possible that a team that started 6 months ago is facing the prospect of being leapfrogged by teams starting last week. This does not mean you can afford to wait – if you want to ship a game, you work with what you have. But it is nonetheless a fascinating phenomenon to live through.

Leave a Comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.