I’m starting an occasional posting of a few items from X and looking at those through the lens of my own work on building a storyworld with AI and 3D.
This article starts out with updates on relevant tools, game engines, and ends with thoughts on a great essay about AI and creativity by the former head of Disney and Dreamworks. If you read only one thing that I mention here, be sure to read that.
Storyworld tools: voice & continuity
AI anime from 松丸 彗吾(keigo matsumaru).
Voice acting example, in Japanese, of the new v4 text-to-speech from ElevenLabs. The music is by Suno; images by GPT Image 2.5; sound effects, editing, animations done with Opus 5.5.
ElevenLabs v4 introduces more capabilities, in more than 90 languages, to direct the emotion of how the text is spoken through audio tags. Lots of examples in the video below.
Google AI: Gemini 3.8 Flash TTS launch.
New voice models also from Google. As good as ElevenLabs? Can I cancel my ElevenLabs subscription? I’m not sure, yet. Matsumaru (see above) ranks ElevenLabs above Gemini for Japanese voice acting.
Prompting guide for Gemini 3.8 TTS: https://ai.google.dev/gemini-api/docs/speech-generation#prompting-guide
Includes a set of tags for non-speech sounds:
Google Research: “Automating coherent long-form video generation”.
From Google Research official account.
In generated video, maintaining consistency in character’s clothing and surrounding environment is improving rapidly. But it’s still difficult as the video gets longer. For details, see the research update in the blog post: Automating coherent long-form video generation. It’s worth a close read even if you’re not using Google tools for video generation.
World Models & Engines
AMD acquires World Labs for $8.2 billion.
I’ve been admiring the work done by Fei-Fei Li and team at World Labs, which was founded in 2024. I didn’t expect them to become part of a chipmaker, but good for them. Fei-Fei Li will now be Executive Vice President and Chief Scientist at AMD.
Storyworlds are a small part of what’s done through the tools developed by World Labs. They’ve been promoting the term spatial intelligence and leaning heavily into robotics and computer vision. On her Substack, Fei-Fei Li describes her vision, which is in a different realm than what we see with OpenAI and Anthropic: the physical world and not just the digital world.
“Without having a focused hardware effort, AI is hobbled in efficiency.”
While my own specific work is in the digital realm, the societal impact of hardware that understands its surrounding physical environment will unleash a range of new consumer products and also new methods of manufacturing.
Images to 3d modeling to storyworlds for games.
Two posts from Anthropic employees caught my attention that showcase the increasing capacity of Opus, Fable, etc., to create game environments.
By GTA, Karpathy is using that as shorthand for an explorable 3D simulation. And that’s exactly what I’m working on.
Alex Albert provided the prompt for the 1906 SF example:
Recreate Market Street, San Francisco as it stood on April 17, 1906, the afternoon before the earthquake, in Blender.
Scope: the Ferry Building up Market to Fifth Street, including the Palace Hotel, the Call Building, the Chronicle Building, Lotta's Fountain and the Emporium.
Before modeling anything, build a source file from: the 1899-1905 Sanborn fire insurance maps (footprints, heights, materials, occupants), the Miles Brothers film "A Trip Down Market Street" (April 1906), period photographs from OpenSFHistory, the Library of Congress and the David Rumsey collection, and USGS topography. Record every building with footprint, height, facade material, occupant, and the source for each fact with a confidence level.
Build everything in Blender Python. No downloaded meshes, textures or HDRIs. Write reusable generators (Victorian commercial facade, mansard roof, bay windows, awnings, painted signage, gas and electric street lamps, cable car, horse-drawn wagon, early automobile) and assemble the street from the source file, so every building traces back to data.
Provide a 10 second video up the street.
Note that it’s a lot more than take this photo and create a game.
It’s really remarkable, though. This would make for a great project for history class. Are digital humanities still a thing?
Later in this article, I point to an essay by Jeffrey Katzenberg who included this quote from Steve Jobs:
“It’s in Apple’s DNA that technology alone is not enough. It’s technology married with the liberal arts, married with the humanities, that yields us the result that makes our hearts sing.”
Sakura Crossing update with Opus 5.5.
I’ve been keeping an eye on the prototyping that GMI Cloud has been doing on this environment. Now they’ve tried it with Opus 5.5.
I also am finding that Opus 5.5 is what I’m reaching for when creating my storyworld.
Thomas Mahler on the difficulty of game dev.
Playing Ori and the Blind Forest is one of the great experiences that my daughter and I shared in her childhood, particularly during the lockdown. After finally escaping the Ginso Tree, (if you know, you know), we threw our arms around each other in a gleeful hug.
Thomas Mahler is the creator of Ori and the Blind Forest, which I think is the best video game from the past decade.
I agree completely, and it’s tiresome to see posts from people saying they’ve one-shot a video game. You can do a lot with these AI systems but a completed, fully functional game is a complex system. You can do a good prototype with Opus, Fable, Astra, but you need a game engine (like Unity, Unreal, Godot) to go further unless you’re game is quite simplistic.
Mahler, who knows a bit about video games, says:
Most of the stuff I see people post online that they proudly state ‘AI one shotted this game!’ are still godawful stuff that nobody would actually consider paying money for. And even the non-one-shots, 99% of those are... well, not great.
You’re not wasting your time learning a game engine.
The game examples we see online are generated in JavaScript, specifically three.js. That’s a great tool but games are pushing the limits. You can generate some cool environments. Here’s a call from an animator about that:
Game engines are intimidating, which is why the web games generated by Opus and Astra are appealing to vibe coders new to game development. But MCP allows you to use your favorite AI to help you learn and build. It’s fine to start with three.js and then go from there into a game engine.
Building with Agents
So much is happening with agents! So much! My head is spinning. I’m in Claude Code all day. This morning, I just started using Grok Bot as well. I still need to get into OpenAI’s Codex, which I’ve only played around with for a bit. But I need to stay focused on what I’m building and not trying out every new technique that comes up.
Coming back around to AI-generated video, here’s a post on a process:
Note that the guy who made the post works for a VC firm and is an investor in some of the tools that he mentions. Is that what a VC does these days? Anyway, there are some good tips in that post.
Show me your prompt.
If you seriously use any of Anthropic’s model, then you need to follow @thariq on X.
I’ve spent the last 3 days working in Claude Code on a single project across 15 sessions. Each session has almost a full context window with sources and iterative prompting. There’s no way that can be just shown to anyone. I do have the key parts hooked into a GitHub repo but it’s not just one prompt. The documentation for the repo is currently more than a hundred files.
Since I’ve been discussing making games with AI, Thariq just made a relevant post as I am writing:
In the replies, he mentions that this is “all prototype quality”. Let’s stress this sentence: “Use AI to work with you to bring your vision to life.”
A wider lens on AI & Creativity
I have a lot more bookmarks from X for this past week, but let’s wrap this up with some thoughts from Jeffrey Katzenberg, the former head of Disney’s animation studio and Dreamworks. He knows a lot about the business of creativity.
That’s a great essay on the matter of AI and Hollywood. It should be required reading for anyone studying the humanities.
AI is not killing creativity. There will still be careers for creative people, but you will need to be a lot more entrepreneurial.
“There will be more seats at the table, and very soon entirely new forms of storytelling.”
I didn’t realize, or had forgotten, that animation in the 1980s was considered a niche form of the business.
“Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process.”
Just prior to that quote, he names some of the most famous filmmakers who have “embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process.”
A quibble I have with Katzenberg’s essay is that he frames it purely in a Northern California (Silicon Valley) - Southern California (i.e., Hollywood) divide. But he was likely writing this for specific people he knew in Hollywood that feared the encroachment of AI on their industry.
Expanded cinema. That phrase, “expanded cinema in the process”. That’s the process now occurring with AI filmmaking. Mix in game development, and we’ll see new forms of cinema and entertainment that will help us enjoy, reflect on, and engage with the world in which we live.
Note on process: I could have had Grok Bot pulled all my X bookmarks and write this post. I did have Grok Bot pull my recent X bookmarks and organize those into topics. But what would be the point of having AI generate this article for me? Sure, that would have been quick, easy, and job done. But it wouldn’t have had any value for me; I wouldn’t have learned anything unless I had sat at my desk this morning for a few hours, read through those posts and thought about how they’re relevant to me. Writing up these notes may be more useful to me than for anyone who reads this. And that is the value of writing for yourself.
























