"A true world model must be a structured, interactive simulator"
Tripo AI’s Simon Song and Yanpei Cao discuss building 3D foundation and world models for next-generation content
Hello and welcome to AI Gamechangers, the newsletter that goes behind the scenes with the people building practical AI tools for the games industry.
This week we’re going deep on 3D. Our guests are Simon Song, founder of Tripo AI, and Dr Yanpei Cao, co-founder of the company. Tripo AI is a platform that’s serving big players in the industry with 3D assets, but has moved well beyond that into world models: generative, interactive environments with persistent state, physics logic and multi-agent support. It’s ambitious stuff, and they talk us through exactly what that means for game developers, UGC (user-generated content) creators and the studios already integrating Tripo AI into their pipelines.
As ever, scroll to the end for a round-up of the latest AI and games news (and remember to browse our archive for complete, free Q&As with leaders in the AI and games space).
Simon Song and Dr Yanpei Cao, Tripo AI

Meet Simon Song, founder and CEO, and Dr Yanpei Cao, co-founder and chief scientist, of Tripo AI.
Tripo AI is the 3D generation platform backed by nearly $200 million in its Series A+ and Series A++ financing rounds, followed by a very recent announcement of $150 million in Series A3 funding, with backers spanning games, automotive, and internet sectors. The most recent round follows a string of model releases, including 8K texture generation and a new take on world modelling with Project Eden.
Simon brings a strong background in AI commercialisation from his time at SenseTime and later became a co-founder of MiniMax. Yanpei brings deep R&D expertise from Tencent’s ARC Lab and a PhD in computer graphics from Tsinghua University. Together they’re building a foundational infrastructure for spatial computing: not just tools for generating assets, but an engine for creating persistent, interactive, multiplayer-ready worlds.
Top takeaways from this conversation:
Tripo AI has raised nearly $350 million this year and counts Riot, EA, Sony, Tencent and NetEase among its partners, with studios integrating AI 3D generation directly into real-time workflows rather than treating it as a standalone tool.
Their priority is eliminating the need for manual asset cleanup entirely, making AI-generated 3D output indistinguishable from human-crafted, production-ready assets in both topology and physical logic.
Their world model initiative takes a different architectural approach to rivals: it maintains a persistent structured “state” of the world, solving for object permanence and reusable, modifiable environments.
They explain that AI won’t replace 3D artists but will act as a force multiplier, shifting creative work away from repetitive modelling and toward stylistic direction, curation and overall experience design.
AI Gamechangers: For readers who haven’t come across Tripo AI yet, can you give us the elevator pitch? What does your platform do, and what problem does it solve for game developers specifically?
Simon Song: Tripo positions itself as the fundamental technological infrastructure and platform powering the world-building engines for next-generation interactive entertainment and global UGC ecosystems.
Unlike traditional AI plug-ins, Tripo is explicitly not designed as a mere efficiency tool for game studio artists. While efficiency tools offer incremental cost reduction, their commercial ceiling is strictly capped. Tripo’s ultimate objective is to provide the infrastructure that blurs the line between players and creators, fundamentally altering the economics of digital content.
“We have always positioned AI as an empowerment tool rather than a replacement. Its value lies in enhancing creative efficiency and expanding the boundaries of what can be created”
Simon Song
The corporate vision is “Empowering every individual to architect interactive, complex worlds with the ease of natural expression.”
The core mandate is radical democratisation. By utilising world models and generative AI, Tripo eliminates technical barriers, enabling users with zero background in 3D modelling or software engineering to seamlessly build rich, interactive experiences.
Positioning a company solely around cost reduction yields diminishing returns. Tripo aims to transform the creative workflow entirely. The goal is to empower ordinary users to orchestrate, simulate, and publish interactive 3D spaces across any accessible platform - making 3D space creation as intuitive and frictionless as uploading a short video today.
Can you give us a quick sense of your own origin story? Do you come from a technical, business, AI or games background?
Simon Song: I completed my undergraduate studies at Johns Hopkins University. Early in my career, I was at SenseTime, where I helped drive several AI projects from the ground up. My primary focus was on the practical application and commercialisation of AIGC (AI-Generated Content) technology within the animation and gaming industries. Later, I also co-founded MiniMax, a company specialising in general-purpose large models.
My academic background has given me a perspective that leans more toward how technology can be truly utilised and integrated into industrial workflows, rather than focusing on the technology in isolation.
Beyond my professional life, I have always been a heavy consumer of content and games. I’ve loved gaming since I was a child; I’m a big fan of series like One Piece, and I enjoy playing board games in my spare time. These interests have had a profound impact on me, fuelling a deep curiosity about the “how” behind content creation. In a way, what I am doing now is the perfect intersection of my lifelong passions and my professional path.
Dr Yanpei Cao: My background sits right at the intersection of AI and computer graphics. But my core drive has always been rooted in content creation.
Growing up, I studied drawing for years but realised I lacked the traditional physical techniques. When I discovered Computer Graphics at university, it clicked: I could use what I was good at (coding and maths) to achieve my passion for art and creation. My ultimate goal has always been figuring out how to make 3D interactive content creation as accessible as possible. I build these tools because I genuinely want to use them myself.
Before co-founding Tripo, I led 3D generative research at Tencent’s ARC Lab and AI Lab. But I also have an entrepreneurial background: earlier in my career, I was the CTO of Owlii, a volumetric capture startup that was later acquired by Kuaishou. That was really where I learned what it takes to build and scale 3D generation systems for the real world.
“When this team first came together, our ultimate vision was always to build the bedrock infrastructure for the next generation of interactive UGC platforms”
Dr Yanpei Cao
Academically, I’ve spent years publishing at venues like SIGGRAPH, CVPR, and NeurIPS. But for me, the goal has never been just about hitting academic milestones. The real driving force behind my work is figuring out how to teach machines to genuinely understand and generate the world. We aren’t just trying to generate static meshes that look good from a certain angle; we are tackling the underlying physics, topology, and interactive structures of 3D environments so that anyone can easily build a playable world.
Now, at Tripo, I direct our R&D across multimodal 3D generation and generative world models. My core focus is taking foundational AI breakthroughs and turning them into industrial-grade production pipelines. I want to completely eliminate the traditional bottlenecks of interactive content creation. Ultimately, my vision is to build the simulation infrastructure for our world, so that any developer or creator can go directly from a concept to a fully rigged, deployable, and interactive environment in a matter of seconds.
You’ve established technical partnerships with some major names in games (Riot, EA, Sony, Tencent, NetEase, among them)! Can you talk us through what those partnerships look like in practice? How are they actively using your AI in their production pipelines today?
Simon Song: Some developers are not viewing Tripo merely as a tool for generating 3D models, but are instead integrating it into real-time content creation workflows. For instance, during our previous collaborative explorations with creative game ecosystems - such as NetEase’s Eggy Party - we observed a very interesting phenomenon.
“The goal is to make AI-generated 3D assets indistinguishable from human-crafted, industrial-grade assets in both structural precision and physical logic”
Simon Song
There are a vast number of creators on the platform designing their own maps and interactive worlds. AI 3D generation allows them to generate props or environmental elements in seconds, which can then be placed directly into the level editor for testing and adjustment. What struck me most at the time was how this approach transformed the rhythm of content production. In the past, these assets typically had to be crafted one by one by professional art teams; now, creators can generate content on demand and use it instantly during the creative process.
From our perspective, this signifies that AI is no longer just an asset generation tool - it is becoming a fundamental part of the UGC world-building workflow.
Tripo AI recently made waves with the announcement of Project Eden, your world model. For the game developers and studios reading this, can you tell us more about what Project Eden actually is, what it brings to the table for next-generation game creation, and how it fundamentally differentiates itself from the other world models?
Dr Yanpei Cao: To understand the architectural shift behind Project Eden, we must first address a fundamental misconception in the AI industry today: confusing a visual hallucination with a world model.
Currently, the industry is largely split into two paths, both of which face severe limitations for interactive game production. On one side, you have “action-conditioned video generation.” While it dominates the current hype, mathematically, it is fundamentally just an autoregressive prediction of 2D pixel trajectories. The state of the world is only implicitly hidden within the context window of recent frames. The moment an object leaves the camera’s frustum, the model effectively forgets it; when the camera pans back, it has to “hallucinate” the object anew. This completely destroys object permanence.
On the other side, you have static 3D scene generation. This provides spatial awareness but strips away the dimensions of time and interactive logic. It gives you a beautifully reconstructed environment, but one that is frozen, with no state transitions or dynamics.
A true world model must be a structured, interactive simulator. It requires a rigorous definition of a State (the objective structural and semantic reality of the environment) and a Transition (how that state evolves over time in response to actions).
With Project Eden, we approached this as architects of a neural-native engine. We completely discarded the monolithic video generation approach and built an architecture that natively decouples State Evolution from Visual Rendering. We architected it across three distinct layers:
The Evolving Structured State: We do not rely on pixels to remember the world. Instead, we maintain a persistent, globally shared world state. To ensure computational efficiency and rigorous temporal deduction, this isn’t a massive dense 4D point cloud; it is a compact, structured/implicit representation. It governs the underlying geometry, object semantics, and the consequences of any action inputs.
The State-to-Observation Interface: When a specific viewpoint is queried, the system extracts deterministic geometric and semantic constraints from that underlying state. Because these representations are derived from a single, unified source of truth, it guarantees absolute spatial alignment and epipolar consistency, no matter how the camera moves.
Generative Rendering: Finally, the generative renderer receives these state constraints and translates them into high-fidelity visuals. The renderer is no longer blindly “guessing” the structure; it focuses entirely on resolving textures, lighting, and high-frequency dynamic details.
By decoupling the underlying logic from the pixel rendering, we have unlocked three system-level capabilities that fundamentally alter game creation:
First, Absolute Object Permanence and Viewpoint Consistency. Because the state is maintained independently of the camera’s frustum, objects do not randomly vanish or mutate. If a player drops an item and returns hours later, the model queries a confirmed objective state rather than relying on a historical pixel context to regenerate it. Long-term memory is natively solved.
Second, Reusable Worlds and Deterministic Control. Traditional video generation is a “one-shot blind box” - the timeline is irreversible. Project Eden allows users and agents to repeatedly intervene, control, and modify an evolving base state. It acts as a reusable, modular sandbox rather than a disposable video clip.
Third, Native Multi-Agent Concurrency. This is the ultimate bottleneck for video models. If you try to support a multiplayer environment using a pure video model, your compute cost scales exponentially with every new perspective added. In our decoupled architecture, the compact underlying state is shared and updated synchronously across all agents. The system only needs to render the multiple views based on individual local coordinates. This makes concurrent, multi-perspective interaction computationally economical and mathematically possible.
Ultimately, Project Eden is not just another asset generator. We are building the foundational infrastructure for spatial computing: a generative interactive runtime where developers can move from pure imagination to a persistent, logical, and multiplayer-ready simulated environment.
Looking at the bigger picture, what does Project Eden mean for ordinary consumers, the broader scientific research community, and the interactive industry as a whole? How does this technology position itself across those different sectors?
Dr Yanpei Cao: That is the core of what makes Project Eden so exciting. We do not view Project Eden simply as a game development tool; rather, we position it as the foundational interactive runtime for spatial computing and embodied AI. Its architecture natively solves the fundamental tension between generative diversity and structural consistency. By doing so, it serves three critical domains simultaneously:
First, for ordinary creators and consumers, Project Eden represents the democratisation of spatial logic. Historically, building a shared, interactive environment required deep expertise in 3D modelling, state machines, and network synchronisation. We are reducing that threshold to near zero. By allowing underlying neural models to handle the complex state transitions and concurrency, regular users can define rules, instantiate environments, and dictate logic entirely through natural language and semantic intent. It shifts the paradigm from passively consuming pre-compiled content to real-time generative interaction, enabling anyone to orchestrate a persistent, multi-agent sandbox.
“Building a true world model required a fundamental architectural shift: decoupling the underlying structured state from the visual rendering. We greenlit Project Eden not as a pivot away from 3D generation, but as a vital, parallel research track”
Dr Yanpei Cao
Second, for the scientific research community (specifically embodied AI): The current bottleneck in training intelligent agents is the simulation environment itself. Traditional simulators rely on rigid, hard-coded assets that lack the infinite diversity of the real world. Conversely, pure video generation models hallucinate, failing at basic collision detection, spatial depth, and object permanence. Project Eden bridges this gap. It provides a neural-native training ground that offers the boundless generative diversity of AI, but strictly constrained by structural and temporal consistency. Researchers can safely deploy agents into highly complex, persistent environments with unpredictable, multi-agent interventions, allowing for robust, closed-loop training and evaluation.
Finally, for the interactive industry, Tripo serves as the foundational engine for accessible interactive content creation. The industry relies on two essential elements: operable 3D assets (objects with valid geometry and physical boundaries) and interactive environments (the physics and rules that govern virtual worlds). Tripo provides both. Our 3D foundation models, including Tripo H3.1 and Tripo P1.0, generate structurally sound, high-fidelity 3D assets, while Project Eden provides persistent, interactive worlds where those assets can exist, interact, and evolve over time. Together, they provide the fundamental building blocks and infrastructure that drive the industry’s shift toward evolvable interactive worlds.
Ultimately, whether it is empowering a casual creator to generate a persistent digital space, providing a logically sound training ground for AI researchers, or allowing studios to scale massive interactive ecosystems, Project Eden provides the architectural infrastructure to make generative, simulated environments a reality.
Given this evolution, how do you define Tripo AI’s core identity today, and where do you sit in the tech ecosystem?
Dr Yanpei Cao: At our core, Tripo is architecting the foundational, neural-native engine for spatial computing and interactive simulation. We are not an application layer or a simple modelling utility; we are the underlying structural bedrock.
From day one, our mission has been to unlock the underlying infrastructure for a universal, interactive UGC ecosystem. So, our trajectory from 3D generative AI to World Models is not a pivot. It is a rigorous, sequential solution to the two fundamental equations of any simulated universe: defining the State and deducing the Transition.
Phase One is establishing the State (asset creation). Traditional interactive development relies on a rigid, pre-compiled pipeline. Artists manually fabricate every mesh and prop, baking them into massive, static databases. With architectures like Tripo P1.0, we proved that production-ready, topologically sound assets can be generated from noise in seconds. This means spatial elements no longer need to be hardcoded or pre-loaded; they can be procedurally summoned into existence precisely when a player’s action or an AI agent’s behaviour demands them.
Phase Two is deducing the Transition (world creation). However, generating standalone assets only provides the static “nouns” of a world; it lacks the “verbs.” A true interactive environment requires a system that comprehends time, spatial depth, kinematics, and physical persistence. It must predict how objects collide, shatter, and evolve, and how multiple agents can interact within that space concurrently. By decoupling the evolving structured state from the generative visual rendering, Project Eden provides the engine that computes these spatiotemporal transitions without breaking the simulation.
Where we sit in the tech ecosystem is precisely at this runtime layer. We are the structural engine that sits between high-level semantic intent, whether from an LLM, a human creator, or an autonomous agent, and the final interactive reality.

We are moving the industry away from the archaic paradigm of loading pre-packaged, linear data structures and shifting it entirely toward a generative interactive runtime. We are providing the foundational infrastructure where the physical rules and the geometry of the world are synthesised, maintained, and interacted with completely in real time.
Can you pull back the curtain a bit on the internal timeline for Project Eden? When did you officially greenlight the development of this world model, and when did you feel the tech had truly matured?
Dr Yanpei Cao: To be perfectly candid, in deep-tech research, you don’t simply sit in a boardroom one day and “greenlight” a world model. There was no sudden pivot or reaction to a market hype cycle. When this team first came together, our ultimate vision was always to build the bedrock infrastructure for the next generation of interactive UGC platforms.
If you look at our timeline, our strategic decisions were based on solving the most critical bottlenecks in interactive content creation. Back in 2023, we focused heavily on 3D generative AI because we recognised its immense, immediate value to the current interactive content industry. We saw a clear opportunity to revolutionise existing production pipelines, and we dove headfirst into solving the algorithmic and representational complexities of spatial generation.
However, generating a standalone 3D asset is only one part of the interactive equation. The other part is “environment and dynamics.”
As we were pushing the boundaries of 3D generation, we observed the broader AI research community attempting to build “world models” primarily through 2D video generation. Because of our deep roots in 3D, we immediately recognised the architectural ceiling of that approach. We knew that relying on 2D pixel trajectories would inevitably fail at object permanence, spatial consistency, and multi-agent interaction.
We realised that building a true world model required a fundamental architectural shift: decoupling the underlying structured state from the visual rendering. And because of our extensive research in 3D generation, we already possessed the exact spatial priors, the geometric intuition, and the massive data curation pipelines required to engineer that decoupled architecture.

Therefore, we greenlit Project Eden not as a pivot away from 3D generation, but as a vital, parallel research track. Today, our 3D generation models are actively transforming industry asset pipelines, while Project Eden tackles the complex algorithms of state transitions and interactive dynamics. For us, these are not competing priorities; they are parallel engines driving toward the exact same endgame.
Earlier this year, you announced a new funding round of $200 million. What does it enable you to do that you couldn’t do before? And what’s the priority for you next on your roadmap? [Note: this interview was conducted just before a further $150 million of funding was announced.]
Simon Song: Our priority is scaling our native 3D representations to handle unprecedented geometric complexity while guaranteeing production-ready topology right out of the box. This funding allows us to scale our models and our proprietary data curation pipelines to a level where we can definitively eradicate the need for manual cleanup or retopology. The goal is to make AI-generated 3D assets indistinguishable from human-crafted, industrial-grade assets in both structural precision and physical logic.
Simultaneously, a significant portion of this capital is dedicated to our World Models initiative, Project Eden. We are actively funding the deep research necessary to solve the most complex bottlenecks in spatial computing: state transitions, physical dynamics, and multi-agent concurrency. We are making the leap from generating static objects to orchestrating dynamic, interactive environments.
There are other companies in the AI-assisted 3D asset creation space. As you all compete for developers’ attention, what makes Tripo AI different, and what do you think it takes to win in this market long term?
Simon Song: I believe the first thing to emphasise is our long-term vision. Since 2023, we have held the very clear conviction that AI 3D generation is not just a short-term tech trend, but will evolve into a fundamental capability across the entire content production workflow. This belief determined our goal from the start: we weren’t just aiming to build a tool that produced “good-looking results”; we focused on its usability in real-world production environments.
“Whether it is empowering a casual creator to generate a persistent digital space or allowing studios to scale massive interactive ecosystems, Project Eden provides the architectural infrastructure to make generative, simulated environments a reality”
Dr Yanpei Cao
On this basis, a core difference between us and some of our peers is that we don’t simply optimise “generation” in isolation. Instead, we design our technology and products around the question: “Can the output actually enter the production pipeline?”
There are many excellent products in the industry today that excel in visual effects or generation speed. However, in actual development, teams care more about whether an asset has a stable structure, clean topology, and whether it can move directly into the production pipeline without requiring extensive post-production cleanup. This is why Tripo has continuously updated and even reconstructed our underlying technical roadmap throughout our iterations.
In the long run, I believe the key to this market won’t be who generates the fastest or whose one-off results are the most stunning. It will be about who can consistently provide “usable” assets. This requires several critical capabilities: output consistency and controllability, support for editable and iterative workflows, and the ability to truly integrate into existing production ecosystems.
Ultimately, AI in this field is more likely to become infrastructure that amplifies a creator’s talent rather than replacing them. The winner in the long term will be whoever best serves the authentic needs of developers and embeds themselves into the production workflow.
There’s been a lot of debate in the games industry about the ethics of AI-generated content. Gamers, for instance, seem to push back against it in the West. How does Tripo AI approach these concerns, especially when working with major studios that might have their own policies on AI use?
Simon Song: From our perspective, we have always positioned AI as an empowerment tool rather than a replacement. Its value lies in enhancing creative efficiency and expanding the boundaries of what can be created, not in diminishing the role of the creator. Especially in the gaming industry, truly high-quality content still relies on an artist’s judgement, stylistic oversight, and creative expression - elements that AI cannot replace in the short term.
“I’ve loved gaming since I was a child; I’m a big fan of series like One Piece, and I enjoy playing board games in my spare time. These interests have had a profound impact on me. In a way, what I am doing now is the perfect intersection of my lifelong passions and my professional path”
Simon Song
In terms of implementation, we maintain a high level of respect for our partners’ strategies and internal standards. When collaborating with major gaming companies, we find they typically have very explicit AI usage policies, covering data compliance, scope of use, and content moderation processes. We align our integration and deployment with these standards to ensure that the technology is utilised within a controlled and transparent framework.
Simultaneously, we are continuously refining our own internal mechanisms. We remain highly prudent regarding data usage, model training, and output management to minimise potential risks as much as possible.
Looking ahead, do you think AI-powered 3D generation will become a standard part of every game studio’s toolkit? Will it be as routine as, say, a game engine? And if so, what does that world look like for the artists and creators who currently make 3D assets by hand?
Simon Song: It’s not just for game studios; I believe AI 3D will become a standard tool across multiple industries. Much like how game engines and DCC (Digital Content Creation) tools have become foundational infrastructure today, AI will likely exist within these tools in the future - becoming a default feature rather than an external tool that requires additional learning or workflow switching.
However, the more significant impact lies in how it reshapes the role of the creator. I don’t believe AI will replace 3D artists; instead, it will act as a force multiplier for their capabilities. In the past, an artist might have spent a vast amount of time on repetitive modelling or creating basic assets. In the future, these tasks will be drastically accelerated, allowing them to focus their energy on stylistic design, creative expression, and overall content direction.
At the same time, the barrier to entry for creation will be lowered, enabling more people to participate in 3D content production. While this may lead to a massive surge in content supply, it also means that aesthetic judgement, original creativity, and the ability to curate the overall user experience will become the truly scarce and valuable resources.
Further down the rabbit hole
Some useful news, views and links to keep you going until next time:
Since conducting this interview, Tripo AI has raised $150 million from games and automotive investors backing its 3D foundation models. Tripo AI now plans to increase investment in world models, with a focus on core algorithm development, data infrastructure, and the recruitment of top global talent. Tripo AI is a sponsor of SIGGRAPH 2026 this month, where Dr Yanpei Cao will deliver a keynote speech on Wednesday, 22 July.
Godot has drawn a line in the sand on AI-generated code. The open-source engine behind Slay the Spire 2 will no longer accept AI-authored contributions, with autonomous agents and vibe coding triggering automatic bans, after a rising tide of slop.
PGC Summit Shanghai lands on 29 July with a dose of AI content in the afternoon, perfectly timed to lead you straight into ChinaJoy, which is running a “Level Up with AI” theme this year. Busy week, if you also factor in CIGDC.
In other event news, the AI and Games Conference returns to London on 10-11 November 2026, a two-day deep dive into how the industry is actually putting AI to work.
Ludo.ai has given its Sprite Generator a workout. The update adds a revamped Pose Generator, new Transition tools for bridging animation clips, and a Rotate Sprite feature that spins fresh viewing angles from a single sprite. The aim, per CTO Jorge Gomes, is saving solo devs and small teams time at the fiddly bits.
Savvy Games Group has signed a memorandum of understanding with Genvid and Massive Studios to provide Saudi studios with access to advanced AI tools, plus training, mentorship, and educational resources. All part of Vision 2030’s games ambitions.
Krafton tested its “co-playable character” at the end of last month. PUBG Ally, its beta AI teammate, is designed to drop into a two-player Arcade mode where you talk to your companion with your voice and she loots, revives and strategises alongside you. It runs on-device via NVIDIA ACE.


