World Models: a trillion dollar opportunity
In the first half of 2026 alone, investors poured $3 billion into the next frontier of AI: World Models. Google DeepMind, Fei-Fei Li’s World Labs, Yann LeCunn’s AMI labs, and Odyssey are all building in this space. In this episode, host Rana el Kaliouby speaks with the co-founder and CEO of Odyssey Oliver Cameron. They dive into the trillion dollar opportunity behind world models, why they are the next frontier of AI, and how they will be applied to a range of industries, from education to healthcare.
About Oliver
- Co-founder & CEO of Odyssey, a leading AI lab for general world models
- Raised $310M Series B at $1.45B valuation in 2026
- Founded self-driving startup Voyage; acquired by Cruise
- Decade pioneering state-of-the-art autonomous vehicles
- Backed by GV, Amazon, NVIDIA, AMD, In-Q-Tel, Jeff Dean, Elad Gil
Table of Contents:
- What world models learn beyond language
- Why world models are drawing massive capital
- Why robots need models of physics
- How self-driving shaped Odyssey's founding
- World models could create trillion markets
- What it takes to train a world model
- Why compute access is now a moat
- How guardrails work for real-time world models
- Episode Takeaways
Transcript:
World Models: a trillion dollar opportunity
Note: Transcripts are automatically generated from episode audio, and are not fully corrected for spelling, grammar, and formatting.
OLIVER CAMERON: Gaming data, in particular for world models, includes so many interesting interactions that these models can learn effectively from. Companies are paying people to play games.
RANA EL KALIOUBY: Don’t tell my son that.
CAMERON: If you think about driving, you can have lots of actual dashcam videos of people driving real cars, but you can also have video game examples, and we don’t want the model to learn from people driving in GTA, right? In world models, I think the company that will win is the company that’s most willing to explore all the applications: driverless cars, robotics, gaming, defense, healthcare, and other industries, even energy.
EL KALIOUBY: In the first half of 2026, investors have poured over $3 billion into the next frontier of AI: world models. Companies like Google DeepMind, NVIDIA, Fei-Fei Li’s World Labs, Yann LeCun’s Ami Labs, and Odyssey are all building in this space. Today, I’m speaking with Oliver Cameron, a co-founder and CEO of Odyssey, on why world models are so hot right now. Oliver and I are diving into the trillion-dollar opportunity behind world models, why it’s the next frontier of AI, and some of the applications that world models unlock, everything from robotics to education and healthcare. I’m Rana el Kaliouby, and this is Pioneers of AI, a podcast taking you behind the scenes of the AI revolution.
[THEME MUSIC]
Oliver, welcome to Pioneers of AI. I’m so excited to have you on the show.
CAMERON: Thank you for having me. It’s great to be here.
EL KALIOUBY: I have to ask you, the name of the company is in the news right now because of the movie Odyssey. Did you take the whole team out to see the movie?
CAMERON: We did. We rented —
EL KALIOUBY: Oh, you did?
CAMERON: A theater, yes.
EL KALIOUBY: Oh my God, that’s awesome.
CAMERON: We actually named the company originally after 2001: A Space Odyssey, this pioneering sci-fi movie about the future and new technologies. So it’s cool, I guess, to see another version of that.
EL KALIOUBY: Yeah. Before we dive in, I want to disclose that I’m an investor in your company, Odyssey, through my fund, Blue Tulip Ventures. In June, you raised $310 million for your Series B round at a $1.45 billion valuation, which is amazing. Congratulations.
CAMERON: Thank you so much.
Copy LinkWhat world models learn beyond language
EL KALIOUBY: Amazing. I want to set the stage. A few years ago, we had the ChatGPT moment, right? We saw large language models become mainstream and ubiquitous, and it’s this idea that you can synthesize large amounts of text, images, maybe voice, even video, to generate stuff, right? And that is kind of what is really magical about large language models. But there’s been an evolution at this frontier, and the next frontier of AI is world models. I would love your definition of world models, and also why now, and why should we care?
CAMERON: Absolutely. I think you nailed it with language models. They really ingest so much language that they deeply understand language and then can simulate language. But importantly, language is a representation of the world that is lossy and biased.
We’ve written things down for thousands of years, but it’s our opinions and our thoughts, and there is so much information that is lost when we transcribe our thoughts or observations into text. So although language models are these incredible intelligences, we think we can go further. To go further, what we really want to train a model to do is understand and simulate the world. And how do you do that? Instead of training it on lots of text, we train it on lots of visual observations of the world. Everything that you could possibly imagine happening in the world is represented as a video that the model is learning from. After it watches enough videos, it deeply understands physics, dynamics, cause and effect, humans, our behaviors, how we work together, and all of these things that you don’t necessarily synthesize in text. So that model, to me, is effectively learning about reality, and I feel that model will just be more powerful and will represent this leap of intelligence that is difficult to comprehend.
EL KALIOUBY: Yeah, I’m actually more excited about the idea of world models and all of the applications they can unlock. I’m more passionate about that than the vision or mission of getting to AGI. I’m curious, how do you think about AGI versus world models? Are they the same in your brain, or are they different?
CAMERON: I think it’s likely that world models play a key role in enabling a superintelligence, and here’s how I think this plays out. One of the applications we’re very excited about with world models is as learning environments for other AIs. To use our own analogy, we live in this world, we live in this universe, and we have all these agents, all these humans, existing within it, and the world is constantly trying to kill us, right?
EL KALIOUBY: Yeah.
CAMERON: Since we became human, we adapt. We learn not to die, and we learn to improve our world, to make it safer, better, and more plentiful. The analogy for a world model is that you have these worlds a world model can provide, these infinite simulations of all these different worlds. Agents can exist and learn within those worlds, and they learn things by being exposed to that world. Interestingly, they also push the world itself to be better. So, to answer your question about superintelligence, it’s very likely that a superintelligence is learning inside a world model.
The world model itself may or may not be a superintelligence, but I think it will play a key role in enabling an intelligence to exist.
Copy LinkWhy world models are drawing massive capital
EL KALIOUBY: Yep. World models are very hot right now. Just in the first half of 2026 alone, venture investors have invested over $3 billion into world model startups. I’ll give a few examples. Obviously, Odyssey raised $310 million. Ami Labs, Siyamak Lahijani’s company, raised $1 billion. Fei-Fei Li’s World Labs raised $1 billion. And even Runway, which is a generative AI video startup, has kind of pivoted to build a world model. Why is there an urgency to build and invest in world models right now?
CAMERON: I think it’s two parts. The first part is that it presents something that feels distinct and really encouraging beyond language models. We could just decide, in the venture community, that language models are it and funnel every dollar into that industry, but I don’t think the venture industry works that way, right? Thankfully. It basically says there are other possible intelligences, other approaches, so let’s fund them, too. And world models, I believe, are one of those. That’s one aspect. I think the second aspect is you just have to look at where the talent is headed. I posted a chart to our Slack today, actually, of the frequency of mention of the term world model in published papers.
You see this incredibly steep rise in 2025 and 2026 in the number of mentions of world models across publications. That shows you that researchers are increasingly spending their time on this thing, world models. So I think the money is following the talent. I also think it dovetails nicely with the excitement in robotics. We’re seeing so much maturation in robotics hardware and —
EL KALIOUBY: More generalized robots, I guess.
CAMERON: Exactly. Humanoids of all different types and other new types of robots. Then it becomes quite clear that those robots need to become intelligent, and how do you do that? World models present a very interesting path.
Copy LinkWhy robots need models of physics
EL KALIOUBY: Can you clarify for us why physical AI, for example, a humanoid robot, needs a world model as opposed to just operating on a large language model?
CAMERON: Yes.
EL KALIOUBY: Let’s imagine we’re a robot here.
CAMERON: So I’ve got a coffee cup here, and we need to grasp the coffee cup, pick up the coffee cup, and place the coffee cup somewhere else. All of those things can be described in text, right? You can type in, “Pick up the coffee cup,” and all sorts of different things. But there is a translation that occurs between the language model and the physical actuation that goes on. In some cases, that’s fine. It just works, and you’d never notice. But our belief is that, in lots of cases, it doesn’t and it won’t, because there is a loss of information that happens between the language model thinking in text and then actually interacting with the physical world. And that can result in task success not being great, meaning it drops the cup, it doesn’t understand from looking at the cup that there is no top on the cup, and that it should be more careful about picking it up, or it doesn’t understand that, with the lighting that’s going on here, this is plastic, not glass or a mug.
So a world model, you can think of as effectively learning the language of the world. It doesn’t speak English. It speaks physics, it speaks dynamics, it speaks lighting, and all of these different things. So theoretically, and still today we’re having to prove this, we believe it will be a better-performing version of this. Other evidence that points to this is that we hear all the time that robotics is bottlenecked on data. If you want your robot to pick up coffee cups all the time, you need to show it lots of examples of coffee cups being picked up. And if you think about trying to make a general-purpose robot where you need to show tons of examples of a thing doing a thing, that really doesn’t feel scalable, right? You’re having to collect lots of examples. Do it here, here, here. You hear from this user, “OK, now more here.” Then all of a sudden you’re scurrying around collecting all these tasks. And there’s been evidence that shows a world model can adapt to a task with dramatically fewer examples necessary, meaning you don’t need to show it 1,000 coffee cup pickups. You just need to show it 10.
EL KALIOUBY: Because it understands the physics involved in doing this action.
CAMERON: Exactly. Another example of this, which is something we’ve been working on at Odyssey and is near and dear to our hearts, is driverless cars. That was my beginning. I spent eight years building driverless cars, and the paradigm of driverless cars, still today, is that you train models on thousands, maybe even tens of thousands, of hours of driving. You show them lots of examples of cars moving, and they thus learn to drive. This is very different from how humans learn, right? When we go to learn at maybe the age of 17, we get behind the wheel for the first time. We don’t sit in front of a television and watch thousands of hours of people driving. What we do instead is say, “I’ve watched my parents drive for a long time. I know which side of the road to drive on. I know that crashing a car is a really bad thing. I know I should be cautious.” So we get to learn to drive very quickly. It takes us less than 30 hours of training to actually be behind the wheel of a two-ton tank, effectively. And we believe that world models can accomplish something similar, where instead of throwing hundreds of thousands of hours at self-driving models, you can instead just show 10 hours of driving demonstrations to a general world model, and it’s already learned all of these concepts of physics, not bumping into things, what side of the road to drive on, and it just applies that knowledge. So long story short, world models should be more adept at interacting with the world.
Copy LinkHow self-driving shaped Odyssey’s founding
EL KALIOUBY: Amazing. I do want to go to your origin story a bit. You and I overlapped a little bit in the automotive industry. When I was running Affectiva, we were building machine learning models that understand human behavior, and one of the applications, it turns out, was in the automotive industry to understand initially driver distraction and drowsiness. But then eventually we expanded that to in-cabin monitoring, and the application was semi-autonomous and fully autonomous vehicles where you want to understand what’s happening inside the vehicle.
Before Odyssey, you started a company called Voyage that was building self-driving cars, and you sold that to Cruise. I would love to hear a little bit about your experience doing that, like what got you into self-driving in the first place, especially given that before that you were at Udacity and in the education space. So what’s the arc here?
CAMERON: Exactly. Really, it’s one guy that I got to work with, a guy called Sebastian Thrun. He was the founder of Udacity, and Udacity taught online concepts like machine learning, robotics, self-driving cars, all these crazy things that were locked up in academia and could now suddenly be taught to everyone on the planet.
EL KALIOUBY: You know, I did a computer vision class for Udacity, which was awesome.
CAMERON: I do remember that, yes.
CAMERON: One day we decided, “Hey, wouldn’t it be cool if we taught people how to build a self-driving car? That seems fun. That seems futuristic. Let’s build a class.” And this was early 2015.
EL KALIOUBY: OK.
CAMERON: So we built this class, and we actually built an open-source self-driving car at the time because we wanted people to have this space to build on. In fact, we wanted students to be able to run code on a car in the real world, which was crazy to think about.
CAMERON: There were safety concerns, let’s say. But we solved those, and it was great. To me, after working in machine learning for lots of different things, it felt like the most magical form of machine learning I’d seen, that we could have machines save lives and interact in the most complex places on the planet. That felt incredible. So in 2016, I started Voyage, my self-driving car startup. Our goal really was to build state-of-the-art driverless vehicles and to find places where the state of the art already worked and have an impact.
So we focused on deployments in places like retirement communities, military bases, these sorts of slower walks of life where you can go from A to B without the crazy traffic of San Francisco. We built that company up over five years. It was an incredible experience, and then I ultimately sold that company to Cruise. I spent two and a half years there scaling driverless systems.
EL KALIOUBY: I’ll be back with more of my conversation with Oliver after this break.
[AD BREAK]
One of my key learnings from starting Affectiva was the importance of getting the timing right.
We were so early. When we were building emotion recognition technology and emotion AI, it was before smartphones existed. Now there are a lot more companies building in this space. When you started Voyage, did you expect that we would now be all in on driverless cars? Did you get the timing right?
CAMERON: I think both yes and no. On the first question, I did think it would proliferate faster.
But I can go outside my house today, call a Waymo, travel 50 miles north fully autonomously, and get back home. It’s amazing. So to the question of Voyage being too early or not, my reflection is that the self-driving industry isn’t one where you can start a company very easily. There haven’t been many, if any, self-driving car startups started recently, right? So it’s kind of hard to say, “Did we pick the right time or not?” I think my lesson from Voyage was that in artificial intelligence, and this is counter to most intuition for other companies, you need to simply do the hardest thing first. In most company building, you try to find a relatively simple thing to do first, and then you build upon that all the way to the most complex thing. My lesson was that the companies that were solving San Francisco in self-driving, like Cruise, were in fact going to get to driverless anywhere faster because they focused initially on San Francisco, not slower. And that was really our whole thesis.
EL KALIOUBY: Huh.
CAMERON: And so that took a lot of —
EL KALIOUBY: Courage.
CAMERON: Courage, yeah, to admit it. But it really felt like that at the tail end. And so I do carry that over into Odyssey.
EL KALIOUBY: I love the advice: Do the hardest thing first. How did your experience at Voyage and Cruise inform your decision to start Odyssey? And as a former entrepreneur, I guess, I’m always curious: Did the idea come about while you were at Cruise, or did you take a sabbatical and then the idea came about? How did that all transpire?
CAMERON: Absolutely. So the technological insight in self-driving was that you had these different models to do different tasks. So you had a perception model to perceive the world. You had a planning model to plan the path through the world. And then you had what we call prediction models. It sounds very generic, I know, but really what a prediction model did was say, “I have a visual observation of the world from the camera. I’m going to predict what the world looks like in five, maybe seven, seconds.”
EL KALIOUBY: Like if it saw somebody walking across the street, it’s going to predict that this person’s going to get in front of the car in the next few seconds, something like that.
CAMERON: Exactly. And that goes for cars turning, merging into your lane — basically, what behaviors are they exhibiting that you can use to predict what they might do next? I always felt that this was magic. The idea that we’re time traveling into the future, albeit five seconds, and that these models tend to be right, that they tend to predict correctly what is going to happen. I felt that was magic. And so really, the origin of Odyssey was that. It was that if you can predict a possible future of the world, that seems really powerful. That should apply to lots of different industries and enable new applications. Let’s go build general, what we now call, world models. So that was the origin.
How it came about, really, was I was at Cruise and met someone who was the first researcher at Wayve, another self-driving car company. We met to go for a ride in a Cruise vehicle around San Francisco, fully driverless at the time, and we just threw around different ideas. That was the one we both came to. A few months went by, we started throwing around more details of the idea, and then we decided to start a company.
Copy LinkWorld models could create trillion markets
EL KALIOUBY: Yeah, that’s amazing. World models have tons of applications. What is the total addressable market? What’s the TAM for world models?
CAMERON: Yeah. So I think it’s easily in the trillions. And I think what you have to do to believe that is basically say that they are going to be integral to robotics, and at some point in the next 10 years, we’re going to see upwards of hundreds of millions, maybe a billion, robots that start to be in our daily lives in many different ways, and that world models will play a key role in that.
I also believe that you can see world models playing more fundamental roles in science and discovering things that will lead to, whether it’s drug discovery or material science, and that world models should take a piece of those discoveries. And then there are industries like education, gaming, health care.
EL KALIOUBY: Can you give us examples? What could the application of a world model look like in the education space?
CAMERON: In education, my belief is that there’s no reason every kid and adult shouldn’t have Einstein in their pocket who is a natural teacher of whatever that person wants to learn. And you might ask, “Well, can’t language models do this today?” And the truth is, they can produce text that sort of resembles Einstein. But a great teacher, from my perspective, is one that’s looking you in the eye, that is able to understand if you’re getting the concept that they’re teaching. If you’re not getting it based on your body language, they change what they’re saying. Then your body language changes, and then hopefully you get the concept, right?
EL KALIOUBY: That is so near and dear to my heart, as you would expect, because I spent years looking at these emotion signals.
CAMERON: Exactly.
EL KALIOUBY: And a world model could do that, right?
CAMERON: Precisely. A world model can have an input of video, and that video is you as you currently are, and it’s continuously generating simulations that are responding to your body language. It simply has the knowledge of body language in pretraining to better react to it and become a better teacher.
EL KALIOUBY: Because world models can unlock so many applications, how are you prioritizing what industries to go after?
CAMERON: Yeah, we have a somewhat unique take on this, I think. In world models, I think the company that will win is the company that’s most willing to explore all the applications.
And I think, again, that kind of goes counter to what lots of early-stage companies do, which is focus, focus, focus on one thing, maybe two things. And we’re quite intentional about saying, “No, we’re going to go very uncomfortably broad.” We’re going to explore driverless cars as much as robotics, as much as gaming, as much as defense, as much as health care, as much as all these other industries, even energy.
That necessitates, I think, building both a culture and a team that is capable of that breadth of exploration. So that’s quite intentional. And I think it all leads back to this is a foundation model, ultimately. It’s a model that can help all these different industries do more, and we need to be the company that has explored the best, most interesting applications.
EL KALIOUBY: I’ve heard you say that we are in the equivalent of the GPT-2 moment in world models. When do you think we’re going to see the ChatGPT moment of world models? And where are we on the spectrum from doing basic research to actually getting closer to commercialization?
CAMERON: My mind has changed on this.
I think this ChatGPT moment will look like one world model that can drive a car, fly a drone, control a robot, generate a game, and teach your kids the ABCs. If one world model can do all of those tasks, then that is a hugely monumental moment, I feel, and that is a GPT-3-esque moment. So that’s what we’re running toward. And it may not feel like the ChatGPT moment. I think that’s difficult. It’s like trying to replicate the iPhone moment in some ways, right? It’s just a relatively unique time. But I think in terms of its importance, what I described — that one model doing all those things — would be more than equivalent in that moment.
EL KALIOUBY: And do you think we’re close? I’m not going to hold you to a time.
CAMERON: I do.
EL KALIOUBY: OK.
CAMERON: I really think it’s very close, and I think in some part it is even closer than I thought.
Copy LinkWhat it takes to train a world model
EL KALIOUBY: I’ll be right back, but first, a quick break. This is the part of the interview that I’m actually most excited about. I would love for you to take us behind the scenes and help us unpack what it takes to build a world model. In my mind, the way I think about it, if you’re building any type of AI, there are three components: the data, the algorithm, and the compute. So let’s start with data. What type of data is needed to train a world model?
CAMERON: So there are a few sources of data. We can think of this as a big pyramid. At the very base of the pyramid is internet video. On the internet, there are boundless volumes of video that describe everything you could possibly imagine about the world, and it’s growing faster than text on the internet, for example, and it’s evergreen. It’s just continuously updating every day. And so that, to us, is the bitter lesson, this term that’s thrown around a lot in artificial intelligence, where you simply go to the best, fastest-growing data source you can for the task that you want. And internet video is the base of the pyramid.
EL KALIOUBY: Can I ask you a question on that?
CAMERON: Sure.
EL KALIOUBY: Because you’ve got YouTube videos of people driving cars and whatnot, but you’ve also got TikTok videos of people on vacation saying, “Hey, this is me,” and whatever.
CAMERON: Right.
EL KALIOUBY: Are these all useful videos, or are you looking for specific types of videos?
CAMERON: I think it’s hard for me, or a human, to say what’s useful to these models. We think of this as a science challenge, meaning you’ve got this very large set of data that you could train with. What filtering should you put in place to ultimately decide on your final data set? What we learn from and find interesting is totally different from what a model does. For example, if you think about the task of driving, you can have lots of actual dashcam videos of people driving real cars, but you can also have GTA or Forza video game examples. We don’t want the model, in the task of driving a real car, to learn from people driving in GTA, right?
EL KALIOUBY: Right.
CAMERON: But we do want it to learn from people driving safely on highways or whatever. A way to think of this is the model needs both the task it’s going to do, but it also needs counterfactuals. It needs to know what not to do. So we simply have to treat it as a science problem, make sure the data set is balanced, make sure it’s got every exposure we need to these things, and then measure the model results to see if it’s accomplishing the tasks we want better and better, find more of that data, and so on.
At the base layer, there’s lots of internet video. Then we go up a layer, and we think of this as expert data, things that we’re going to acquire from different data companies. One layer of that is robotics data. We need manipulation data. We need lots of different form factors of robots, things like that. Then the next layer up is gaming data. We very much believe that gaming data encodes lots of interesting human behaviors and human interactions. At the very top, that might look like literal video that is incredibly targeted at the specific application. Maybe, I don’t know, I’m making this up right now, but you need a doctor talking about a specific topic, so you’d literally go and find a company that will pay a doctor to say those things on camera, and then that goes at the very top of your pyramid.
You now have this data set that really, if you were to lay it out somehow and zoom out, you would say everything about the world is included here. Every observation you would ever want to know about how physics works, about how the world works, is right there in front of you. And then you train the model on that data set. Maybe it’s a good time to talk about the architecture. These models today, the state of the art, are what’s known as autoregressive diffusion transformers, or ARDITs. These models are autoregressive like a language model. They can’t see the future. All they can see is the past and the current state to then predict the next state. It turns out diffusion is this extraordinary way of doing that, to represent really realistic pixels. And then transformers are just this very general learner, right? They can represent concepts in incredible depth and learn from data of many types. So you bundle all of that together, you have an ARDIT, an autoregressive diffusion transformer. You train it through multiple steps of training, or phases of training.
EL KALIOUBY: Yeah. I am seeing kind of a pattern where world models and physical AI are spurring a whole economy where companies are literally strapping cameras on hotel workers, right? If you want to train a robot to fold laundry, then you want lots of examples of a human doing that. So are you partnered with companies doing that? Do you have a team going around with cameras collecting data? How are you doing this?
CAMERON: You mentioned this topic earlier: how do you know if you’re too early? I think this is a great example of where it feels like we’re in the right time and place, because there is this series of tailwinds moving us faster. I mentioned at the very beginning that the volume of world model research is spiking like this. Internet video growing faster and faster is another. And then this is another, which is that there are so many marketplaces now that have sprung up in the last 12 months that will find you and sell you any data you could possibly imagine, and they’ll do it really fast.
EL KALIOUBY: Give us examples.
CAMERON: You mentioned some of the best ones. The idea that we could have literal insight into factory lines in lots of countries around the world, of humans putting things together, disassembling things, assembling things, and that we could have data sets that represent just that, is incredible. It goes to cleaning tasks, or stocking shelves, or stacking shelves, things like that. There are just tons of examples of these marketplaces solving data gaps that exist. Another that I think has become increasingly prominent is gaming.
It just so turns out that gaming data, in particular for world models, includes so many interesting interactions that these models can learn effectively from. So what you have now are companies that are paying people to play games.
EL KALIOUBY: Oh.
CAMERON: It will become an industry. It already has. So I think if those companies didn’t exist, and we had to do all of that ourselves, that is a major investment we’d have to make that now we don’t.
Copy LinkWhy compute access is now a moat
EL KALIOUBY: Fascinating. OK, let’s talk about compute.
CAMERON: Yes.
EL KALIOUBY: Your thesis, actually, is that access to compute is a key competitive advantage as you’re building these world models. Say more about that.
CAMERON: Yes. This is also where my mind has shifted over time. I think if you were to go back six months ago, the availability of compute was there, and the alarm bells weren’t ringing. There were still shortages, but if you knew where to go, you’d get your hands on it.
EL KALIOUBY: And world models are obviously very compute-intensive, right? Both in training and inference.
CAMERON: Exactly, yes. In particular, inference is much more intensive today, at least. And then a thing happened, which was OpenAI raised 120 billion dollars, and Anthropic raised probably very similar amounts in aggregate. And then all of a sudden, the alarm bells started ringing that these companies were simply buying up every possible amount of compute they could from every provider, and deals that could be done are now not happening.
In fact, during our fundraise, we treated this as the problem to solve. So in this fundraise, we welcomed AMD and Amazon as investors, and we welcomed NVIDIA, and we signed pretty significant compute deals as a part of this fundraise to ensure that we’re not held back by any means on compute. I think of it absolutely as a strategic advantage to have access to the compute that we do, and it’s important we don’t rest on that.
EL KALIOUBY: Yeah. I would love for you to comment: with all these other world model companies building in the space, how is Odyssey’s approach different?
CAMERON: Yes. I see our main competitor base as almost all doing different things. I think maybe over time we’ll converge as an industry to a singular approach, but not right now. Our belief is that the biggest company that emerges from this field will have truly built a foundation world model. It is a model that can solve all of these different tasks, all of these different applications, across a vast variety of industries. So what you learn, meaning the data set, how the model works, and what it can output, are going to influence that substantially. I don’t see any other approach that is as foundational as our approach, in that it’s going to enable all these markets, all of these applications. How do you build the most foundational technology, the most powerful technology? We believe our approach is the path to that.
EL KALIOUBY: Yeah. Back to doing the hard things first, it sounds like you’re just doing the hard thing.
CAMERON: Yes.
Copy LinkHow guardrails work for real-time world models
EL KALIOUBY: So I want to talk a bit about guardrails. There’s a lot of lessons learned from the large language model space around hallucination and how we can guard against all of that. Is there a similar kind of framework for thinking about guardrails in the world model space?
CAMERON: There is. I think it’s less mature, but this is absolutely essential because these systems are going to be operating in the physical world, and that demands much more rigor. The guardrails are tough to get right because of the real-time nature of these models, right? These are models that are operating in real time, and language models have the benefit of extra latency if need be. World models don’t get that benefit. When you’ve got 50 milliseconds to simulate forward in time, these safeguards become a real research challenge. So it is an open problem. We’ve made some really great progress on many different topics within safeguarding. One of the areas where I’m seeing the most encouraging progress is harnesses for world models. This has obviously become a real performance improvement in language models.
EL KALIOUBY: What do you mean by that?
CAMERON: For example, in language models, the harness is Claude Code, and the model is Claude. The harness really gives the model focus. It says, “Do this thing. Here’s my loop to make you do this thing.” It turns out that when you give the model a little sense of rules, it can perform better in those applications. For example, in world models, a harness might be for driverless cars, and that harness might say, “Okay, world model, you’re driving a car. You need to adhere to the road rules. Here is the side of the road to drive on. I’m also going to help you through different events that you might see. And don’t hallucinate, right? Just make sure you deal with the world as it is,” all those kinds of things. I think safeguarding will happen in multiple ways, one of which is at the harness level. The other is at the model layer itself, and we’re investigating and exploring both.
EL KALIOUBY: Very cool. All right, my last question: What are you most excited about for the next milestone at Odyssey? And what do you envision world models will unlock over the next five years?
CAMERON: The most exciting thing internally at Odyssey today is that single world model that can do many of those virtual and physical tasks. If a world model can drive a car, control a robot, fly a drone, generate a game, and even play a game, I think that is a landmark moment and something that demonstrates the power of these models very viscerally. I’m very excited about that. Over the longer term, I very much believe that a model learning physics to the depth these models are learning will uncover things about reality that we cannot comprehend today. It just doesn’t seem unreasonable to me that a model observing physics and reality to the incredible extent that these models are will make scientific breakthroughs.
EL KALIOUBY: Amazing.
CAMERON: And that those breakthroughs will lead to amazing things. So that, on the longer time horizon, is what I’m very excited about too.
EL KALIOUBY: Very cool. That’s a great way to end our conversation. Oliver, thank you so much for joining us on Pioneers of AI.
CAMERON: Thank you, Rana. This was great.
Episode Takeaways
- Pioneers of AI host Rana el Kaliouby sits down with Odyssey CEO Oliver Cameron to explain why world models, trained on video instead of text, may be AI’s next big leap.
- Oliver Cameron argues that robots and self-driving cars need models that understand physics, lighting, and cause and effect, not just language, to act reliably in the real world.
- Drawing on his Voyage and Cruise experience, Oliver says Odyssey grew from the insight that predicting the next few seconds of the world could power far more than autonomy.
- He makes the case that world models could unlock trillion-dollar markets across robotics, education, health care, gaming, and science, especially if one model can span many tasks.
- Oliver also pulls back the curtain on the race itself, from internet video and paid gaming data to scarce compute and new guardrails for real-time AI acting in the physical world.