Course
Imagine you ask a friend to pass you an apple from a fruit bowl. They look at the bowl, work out which one is the apple, and pick it up gently enough that it does not get squashed before passing it to you.
Now let’s ask a chatbot to do the same thing. It can tell you exactly how to pick up an apple, which fingers to use, and roughly how hard to squeeze, but it cannot actually pick one up. It has no hands and no idea what an apple feels like.
Embodied AI is the field trying to close this gap. In simple terms, it is AI that has a physical body (like a robot, a car, or a drone) and learns by doing things in the real world rather than only reading about them. You will also see the closely related term “physical AI” used for the same idea.
Right now, the field is having a big moment. At CES in January 2026, NVIDIA’s Jensen Huang called it “the ChatGPT moment for robotics.”
Money has followed. Dealroom data reported by CNBC shows robotics companies raised $55.8 billion in 2026 by early June, nearly double the previous annual record.
In this article, I’ll cover:
- What embodied AI is and how it compares to LLMs and classical robotics
- Why physical AI is so much harder than digital AI
- How embodied AI works through the perception-reasoning-action loop
- How robots are trained in simulation, and what the sim-to-real gap is
- Where embodied AI is being used in 2026, and which companies to watch
- What all of this means for data scientists
Don’t worry if you have never worked with a robot before. Some familiarity with machine learning will help, but I will build every idea from the ground up so you can follow along either way.
Embodied AI In a Nutshell
- Embodied AI is AI built into a physical body, such as a robot, car, or drone, that learns by acting in the world rather than only from text and images.
- Its main bottleneck is data: Open X-Embodiment, one of the largest open robotics datasets, holds just over 1 million real trajectories, versus trillions of tokens for LLMs.
- Teams work around this with simulation, domain randomization, and pretrained vision-language backbones.
- In 2026, it is in early production in warehouses and robotaxis, while most humanoid deployments are still narrow pilots.
What Is Embodied AI?
Embodied AI is an artificial intelligence system built into a physical body, such as a robot, an autonomous vehicle, or a drone, that perceives its environment through sensors, reasons about what to do, and acts on the world through motors, learning from the outcomes of its own actions.
The key term here is embodied.
Most AI models learn by observing data that someone else has collected. An embodied system learns by interacting with its environment, such as pushing a cup, watching the cup slide, and updating what it knows about how cups behave when pushed.
Here is an example. A vision model trained on millions of photos of doors knows what a door looks like. A robot that has actually tried to open 500 doors knows that some doors push, some pull, some are heavy, and some have handles you need to press down first.
You cannot get this depth of knowledge from a photo.
I like to put it this way: most AI learns about the world by reading about it; embodied AI learns about the world by touching and experiencing it.
This one difference changes the data you need, how you train, and how things can go wrong.
Embodied AI vs. large language models vs. traditional robotics
So far, I have compared embodied AI with a chatbot. Let’s now compare it with the 2 things people most often confuse it with: large language models (LLMs) and traditional rule-based robotics.
| Embodied AI | Large language models (LLMs) | Traditional rule-based robotics | |
|---|---|---|---|
| Learns from the environment | Yes, through interaction and feedback | Mostly no, learns from a fixed text corpus | No, behavior is programmed by hand |
| Takes physical actions | Yes | No | Yes |
| Trains on internet data | Partly (its vision and language backbones) | Yes, trillions of tokens | No |
| Generalizes to new environments | Moderately, and improving quickly | Well, within language tasks | Poorly, breaks when the setup changes |
| Commercial maturity in 2026 | Early production (warehouses, factory pilots) | Mature and widely deployed | Mature, decades of factory use |
The last 2 rows are the ones I want to highlight.
Rule-based industrial arms have welded cars for decades, and they are extremely reliable, but only because nothing inside their work cell ever changes. Embodied AI tries to keep that reliability while letting the world be messy, which researchers call generalization.
LLMs and embodied AI both learn from huge amounts of data, but they solve very different problems. An LLM takes text in and predicts which token should come next, whereas an embodied system takes camera frames, joint angles, and touch readings in and predicts which movement should come next.
In short, LLMs know about the physical world, while embodied AI acts in it. Back to our apple: an LLM can describe exactly how to pick one up, and an embodied AI robot can actually do it.
The cost of a mistake is also very different. If an LLM gets something wrong, you get a slightly odd sentence. If an embodied system gets something wrong, a glass can end up on the floor, or something worse.
Now here is the part I think a lot of people miss: the two work together.
In vision-language-action (VLA) models, a pretrained vision-language model sits at the core of the system. It reads the camera images and your instruction (“put the fruit in the bowl”), then passes its understanding to an action decoder that produces the actual motor commands.

How a pretrained vision-language model connects to robot actions in π₀ (pi-zero). Image from source: Black et al., 2024
Why Is Physical AI So Much Harder Than Digital AI?
Physical AI is harder than digital AI mainly because of data.
Language models advanced rapidly thanks to decades of people sharing text, code, and images online, but there is no comparable source of real-world sensor data. There are far fewer streams of joint angles, force readings, and camera frames from robots interacting with everyday objects.
Let’s look at the data requirements in more detail. An embodied system needs 3 types of data that are hard to find:
- Sensor data recorded from the robot’s own point of view, from cameras, depth sensors, lidar, and touch sensors
- Object physics, such as how things slide, bend, tip over, and break
- Action outcomes, meaning a record of what the robot did and what happened afterward
Now let’s put some numbers on this.
Frontier LLMs are trained on trillions of tokens. Open X-Embodiment, one of the largest open robotics datasets, pooled data from 22 robot types across 21 institutions and ended up with just over 1 million real trajectories.

The robots and tasks pooled in the Open X-Embodiment dataset. Image from source: Open X-Embodiment Collaboration, 2023
Why so little? Every trajectory needs a real robot doing a real task, usually controlled by a human through teleoperation, which is slow and expensive.
This gap is also why physical AI generalizes more slowly than digital AI. A model handles variation best when it has seen something similar before, and a million trajectories cover far fewer kitchens, lighting setups, and object shapes than trillions of tokens cover sentences.
I would keep this data problem in the back of your mind for the rest of the article. Many of the design decisions you will see below, from simulation to domain randomization to borrowing pretrained vision-language backbones, are workarounds for it.
How Does Embodied AI Work? The Perception-Reasoning-Action Loop
Embodied AI works by running a continuous loop of 3 stages:
- Perception: sense the world
- Reasoning: decide what to do
- Action: move the body
This loop repeats many times per second, and every movement changes what the robot senses next, so the output of one cycle becomes the input to the next one.
The same loop sits behind world models. Ha and Schmidhuber’s 2018 “World Models” paper split an agent into a vision module, a memory module, and a controller, and tested it on a car in a racing game. Embodied AI applies the same idea to a full robot.
A good way to understand the loop is to imagine catching a ball:
- First, your eyes track the ball and work out where it is and how fast it is coming.
- Next, your brain predicts where the ball will be in half a second and decides where your hand should go.
- Finally, your arm moves, and as the ball gets closer, you keep adjusting.
A robot does the same thing, just with cameras instead of eyes and motors instead of muscles.
Let’s take each stage in turn.
Perception: making sense of the physical world
Perception is the stage that turns raw sensor readings into something the rest of the system can reason about. A single camera is rarely enough, so most robots combine several sensors:
- RGB cameras for color and appearance; robots usually carry several to capture the layout of the scene
- Depth sensors and lidar for distance and 3D shape
- Proprioception, which is the robot’s sense of its own body (joint angles, velocities, and torques)
- Tactile sensors in the fingertips for contact, pressure, and slip

Examples of perception outputs from Gemini Robotics-ER 1.5. Image from source: Google DeepMind
Perception models then turn these readings into spatial maps, object identities, obstacle positions, and a general understanding of the scene.
Most of this is standard computer vision, such as object detection, segmentation, and depth estimation, usually built on pretrained vision transformers like DINOv2 or SigLIP.
In my opinion, though, touch is the most underrated part of perception right now. A camera can tell you there is a glass on the table, but touch is what tells you the glass is starting to slip out of your grip.
To give you an idea of where things are, XELA Robotics showed a robotic fingertip at Automate 2026 with 30 three-axis force-sensing points packed into the pad. The company also says its uSkin sensors can detect forces as small as 0.1 gram-force.

A uSkin robotic fingertip containing an array of three-axis tactile sensing points. Image from source: XELA Robotics
Reasoning: world models and planning
Reasoning is the stage that decides what to do next. In newer systems, it usually has 2 layers:
- World model (lower layer): an internal representation that predicts how the environment will respond to an action before the robot takes it
- Planning (upper layer): breaks a big goal into smaller steps
World models are what let a robot imagine before it acts.
DreamerV3 is a good example. It learns a compact model of its environment and then trains its policy almost entirely inside that imagined world.
I use DreamerV3 in my own research, and what has surprised me the most is how much more time the agent spends practicing in its “head” than in the actual environment. That matters because training becomes faster and much cheaper.

DreamerV3 trains its policy on imagined rollouts from its world model. Image from source: Hafner et al., 2023
Planning, on the other hand, is increasingly handled by language models.
A good example is Google DeepMind’s Gemini Robotics-ER 1.5, which acts as a high-level planner. Given a task like sorting waste according to local recycling rules, it can look up the rules with Google Search, break the job into steps, and hand each step to the Gemini Robotics 1.5 VLA model that does the physical work.
You can think of the language model as the manager and the VLA model as the worker who actually does the job.

Gemini Robotics-ER 1.5 plans the task and delegates physical execution to the Gemini Robotics 1.5 VLA. Image from source: Google DeepMind, 2025
Action: turning decisions into movement
Action is the stage that converts a decision into precise commands for every joint and motor, usually tens to hundreds of times per second (measured in hertz, or Hz). The low-level part of this stage is called control.
For example, a command such as “Move the gripper 5 cm to the left” sounds simple, but on a 7-joint arm it has to become a specific torque for each joint motor, recalculated continuously as the arm moves.
The hard part is that reality rarely matches what the robot expected. Objects slip, floors are slightly uneven, the lighting changes when someone opens a blind, and a cable bends differently than predicted.
A good controller notices the difference between what it expected and what actually happened, then corrects straight away.
I think action is the hardest stage to get right in messy, unstructured environments. A perception mistake can be fixed by looking again and a planning mistake can be fixed by replanning, but a control mistake happens in real time.
How Is Embodied AI Trained? The Role of Simulation
Embodied AI is trained on a mix of simulation and real-world data. Simulation means virtual environments filled with physics-accurate 3D objects, where a robot can practice millions of times before it ever touches real hardware.
Remember the data problem from earlier? Simulation is one of the main workarounds for it, especially for reinforcement learning of locomotion and manipulation. Foundation models such as π₀ still lean heavily on real demonstrations, so most teams use both.
Collecting data in the real world has 3 big problems:
- Slow: one robot can only collect a few hundred demonstrations a day
- Expensive: robots break, and humans need to supervise them
- Dangerous: especially with heavy humanoids or cars
A simulator addresses all 3 at once. You can run thousands of virtual robots in parallel on a single GPU and let them fail and restart as often as needed. This condenses years of robot experience into a few hours.
What makes a good simulation environment
Not every simulator is equally useful for training robots. From working with robot arms in simulation myself, I would say a good one needs 3 things:
- Physical accuracy, so contact, friction, and deformation behave as they do in the real world
- Object diversity, so the robot sees thousands of different mugs instead of the same mug a thousand times
- Domain randomization, so lighting, textures, masses, and friction change between training episodes (more on this in a moment)
The 2 platforms you will run into most often are NVIDIA Isaac Sim (with Isaac Lab) and MuJoCo:
| Platform | Developer | What it is good at | Best for |
|---|---|---|---|
| NVIDIA Isaac Sim and Isaac Lab | NVIDIA | Photorealistic rendering, sensor simulation, thousands of parallel environments on GPU | Large reinforcement learning runs and synthetic camera data |
| MuJoCo | Google DeepMind (open source) | Fast, accurate contact physics, lightweight, huge research community | Research, locomotion and manipulation benchmarks |
The newest addition is Newton, an open-source, GPU-accelerated physics engine built by NVIDIA, Google DeepMind, and Disney Research and managed through the Linux Foundation.
Newton 1.0 was released at NVIDIA GTC in March 2026. It plugs into Isaac Lab as a physics backend, with MuJoCo Warp as one of its solvers and extra solvers for deformable objects like cables and cloth.
Deformables matter more than they sound. Older simulators handled cables and fabric badly, and factories need robots that can handle both.
The sim-to-real gap
The sim-to-real gap is the drop in performance that happens when a policy trained in simulation is moved onto a real robot, and it is one of the biggest problems in embodied systems.
For example, simulated friction is never quite right, simulated cameras are too clean, and simulated objects can be too perfect, so a policy that succeeds 95% of the time in simulation can break the moment it meets real hardware.
There are 2 main ways teams close this gap:
Domain randomization during training
Rather than trying to make the simulator perfect, teams randomize it in many different ways.
Tobin et al. (2017) randomized textures and lighting so heavily that the real world looked like just another variation to the model. OpenAI’s Rubik’s Cube robot hand went further by automatically widening the randomization as the policy got better.

Fine-tuning on real data after simulation
A policy pretrained in simulation often needs only a small set of real demonstrations to adapt, because it has already learned the general shape of the task.
Neither technique removes the gap completely, so teams keep working to reduce its impact on real-world performance.
Embodied AI Examples in 2026
Embodied AI now runs outside the lab in several industries, though very unevenly.
The pattern I see is that it lands first wherever the environment changes too much to script by hand, but where a human can still step in if something goes wrong.
Autonomous vehicles
Self-driving cars are one of the most data-hungry kinds of embodied AI.
The Waymo Driver has traveled nearly 200 million fully autonomous miles, and it also trains and tests on billions of miles in simulation, according to Waymo’s February 2026 post on its World Model. Simulation lets Waymo replay real incidents and generate rare scenarios (for example, a child running out between parked cars) that a test fleet might never meet on real roads.
Self-driving cars are also a good example of the perception-reasoning-action loop running at full speed. A Waymo vehicle perceives with cameras, lidar, and radar, reasons about what every pedestrian and vehicle is likely to do next, and acts by controlling the steering and brakes many times per second.
Warehouses and logistics
In my opinion, warehouses are the most natural commercial use of embodied AI right now.
The shift is from fixed-path robots that follow lines on the floor to systems that can deal with dynamic, messy inventory, such as picking from a tote with 40 different products thrown in together.
Agility Robotics’ Digit humanoid has been moving totes at GXO Logistics under a multi-year agreement signed in 2024, and Boston Dynamics’ Stretch robot unloads boxes from trucks.

Agility Robotics’ Digit moving totes between autonomous mobile robots and a conveyor at a GXO warehouse. Image from source: Agility Robotics/The Robot Report
Healthcare robotics
Healthcare is where I expect embodied AI to move the most carefully.
Surgical systems such as Intuitive’s da Vinci are already used in hospitals worldwide, but a surgeon controls them, and most AI research here focuses on assistance, such as instrument tracking, camera control, and suturing support.
In healthcare, regulatory approval is often a bigger constraint than model capability, and rightly so. I would personally expect assistive and training uses to arrive long before anything close to autonomous surgery.
Humanoid robots on factory floors
Humanoids get the most attention and are among the least proven.
At CES 2026, Boston Dynamics unveiled the production version of Atlas and said every unit it builds in 2026 is already committed to Hyundai and Google DeepMind. Figure, meanwhile, ran an 11-month deployment of its Figure 02 humanoid at BMW’s plant in Spartanburg, South Carolina, loading sheet metal.
I want to be honest here, though: most humanoid deployments are still pilots doing a narrow set of tasks. Hyundai, for example, plans to start using Atlas for parts sequencing in 2028, a sign of how carefully even the biggest backers are moving.
Key Embodied AI Companies to Know in 2026
Now let’s talk about who is actually building all of this. Here are the 5 organizations I think anyone new to embodied AI should recognize:
| Organization | What they build | Why it matters |
|---|---|---|
| NVIDIA | Isaac Sim, Isaac Lab, Newton, GR00T, and Cosmos | The simulation and compute stack most teams train on |
| Physical Intelligence | π₀ family of robot foundation models | General-purpose robot policies with open checkpoints |
| Agility Robotics | Digit humanoid | Built for warehouse logistics and already working with customers |
| Boston Dynamics | Atlas, Stretch, and Spot | Moving Atlas from research demos to production units in 2026 |
| AMI Labs (Yann LeCun) | JEPA-style world models | A contrarian bet on world models that do not generate pixels |
NVIDIA
NVIDIA mostly builds the tools around the robot rather than the robot itself.
Isaac Sim handles photorealistic simulation, Isaac Lab handles reinforcement learning on top of it, and Newton can now provide the physics. Many robotics teams train on NVIDIA software even though NVIDIA does not sell a general-purpose robot.
Physical Intelligence
Physical Intelligence built π₀, a 3.3-billion-parameter vision-language-action model that uses flow matching to produce smooth chunks of actions for tasks like folding laundry.
It is the best example I know of a general-purpose robot model you can actually download, through the openpi repository.
If the concept of a “foundation model” is new to you, our Introduction to Foundation Models guide explains the idea well.
Agility Robotics
Agility Robotics took the narrow route with Digit.
It was designed around moving totes in warehouses from day one, and that narrow focus is a big reason it reached paying customers earlier than most humanoids.
Boston Dynamics
Boston Dynamics took the opposite route with Atlas.
It spent years pushing the robot’s athletic limits (you have probably seen the parkour videos) before turning it into an all-electric factory product that Google DeepMind plans to power with its Gemini Robotics models.
AMI Labs (Yann LeCun)
AMI Labs is Yann LeCun’s Paris-based startup, which closed a $1.03 billion seed round in March 2026 at a $3.5 billion pre-money valuation. The round was co-led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital, and Bezos Expeditions, with NVIDIA and Samsung among the other investors.
LeCun has argued for years that predicting every pixel of the future is wasteful, so his Joint Embedding Predictive Architecture (JEPA) makes its predictions in an abstract representation space instead, an idea Meta already tested with V-JEPA 2.
What makes AMI worth watching is that it goes against where most of the field is heading. It has not shipped anything yet, though, so I would treat it as a research bet with a lot of money behind it rather than a proven approach for now.
Embodied AI for Data Scientists: Where Do You Fit In?
Embodied AI is not only a robotics engineering problem. A lot of the work, and I would argue most of it, is machine learning work.
The skills that carry over most directly are:
- Computer vision, for perception models and scene understanding
- Reinforcement learning (RL), for training policies in simulation (if you have come across RLHF when reading about LLMs, it belongs to the same family of ideas)
- Simulation and synthetic data, for creating training data that you cannot collect in the real world
- Model evaluation, because measuring whether a robot “succeeded” is much messier than measuring accuracy on a test set
- Data pipeline engineering, for keeping video, joint states, and actions in sync across thousands of episodes
Try the perception-reasoning-action loop yourself
You do not need a robot to get started. The gymnasium library ships with MuJoCo environments, so after running pip install "gymnasium[mujoco]" you can run the whole loop on a simulated robot:
import gymnasium as gym
# HalfCheetah is a 2D cheetah-like robot simulated with the MuJoCo physics engine
env = gym.make("HalfCheetah-v5")
observation, info = env.reset(seed=42)
for step in range(1000):
# Perception: the observation holds joint angles and velocities
# Reasoning: a trained policy would choose the action here, we pick randomly
action = env.action_space.sample()
# Action: apply torques to the joints and move the physics forward one step
observation, reward, terminated, truncated, info = env.step(action)
if terminated or truncated:
observation, info = env.reset()
env.close()
Right now, the cheetah just flails around because the “reasoning” step is env.action_space.sample(), which picks a random action every time.
Replacing that one line with a trained policy is exactly what reinforcement learning is for, and I would recommend trying it before moving on to anything bigger.
What Are the Limitations of Embodied AI?
Embodied AI is moving fast, but it still struggles in 4 places, and I think it is worth knowing them:
- Recovering from mistakes. Most training data shows successful attempts, so robots have seen very few examples of what to do when a grasp slips halfway through a task.
- Long tasks. Small errors add up when you chain steps together. If each step succeeds 95% of the time and failures are independent, a robot only finishes a 20-step task about 36% of the time.
- Speed on the robot itself. Running a model with billions of parameters fast enough for real-time control on the robot’s own hardware is still hard, which is why many systems split slow reasoning from fast control, sometimes running the slow part off the robot.
- Unfamiliar environments. Robots handle scenes similar to their training data well and noticeably worse outside it, which takes us straight back to the data problem from the start of this article.
Final Thoughts
Embodied AI gives machines a physical body and the intelligence to use it. It works through the perception-reasoning-action loop, is held back mainly by how little physical data exists, and in 2026 is moving out of research labs and into warehouses and the first factory pilots.
What really surprises me is how the question has shifted. A few years ago, people asked whether LLMs would make robotics research irrelevant. Instead, LLMs are becoming the reasoning layer inside embodied systems, and the 2 fields are growing into each other.
Remember the apple from the start of this article? Knowing how to pick it up and actually picking it up turned out to be 2 very different problems, and embodied AI is the field working on the second one.
If you want to build the ML foundations for this, start with our Reinforcement Learning in Python track. For the language side of the stack, our Large Language Models (LLMs) Concepts course and our Developing Large Language Models track are good next steps.
Embodied AI FAQs
Is embodied AI the same as robotics?
Almost! Robotics is the larger field of building and controlling physical machines whereas Embodied AI specifically refers to robots that learn from interacting with their environment rather than following hand-coded rules.
Do I need a physical robot to start learning embodied AI?
No. Simulators like MuJoCo and NVIDIA Isaac Sim lets you train and test policies entirely in software, which is how most research and industry teams work.
What programming languages and tools are used in embodied AI?
Python dominates, paired with frameworks like PyTorch or JAX for model training, plus simulation tools like MuJoCo, Isaac Lab and reinforcement learning libraries like Gymnasium or Stable-Baselines3.
How is embodied AI different from computer vision?
Computer vision is a component of embodied AI. A vision model can classify, segment or detect objects in an image but an embodied system must also decide what to do about what it sees and then physically act on it.
Can embodied AI models be fine-tuned like LLMs?
Yes. Just as you can fine-tune an LLM on your own text data, robot foundation models like pi0 can be fine-tuned on a small set of task specific demonstrations, letting you adapt a general-purpose policy to a new robot or task without training from scratch.



