← Back to all posts

PI-WorldModel: Teaching AI How to Dream Without Hallucinating Physics

machine-learningworld-modelsroboticsembedded-systemscpytorch

PI-WorldModel: Teaching AI How to Dream Without Hallucinating Physics

Why standard neural simulators imagine impossible motion, and how embedding basic motion rules creates a tiny, ultra-fast model that runs on microcontrollers.

I have open-sourced PI-WorldModel on GitHub: an action-conditioned latent world model designed to eliminate physics hallucinations in AI. Weighing just 14,025 parameters (~55 KB), it replaces guesswork with basic laws of motion, running on microcontrollers in pure C at 2.94 microseconds per step.

PI-WorldModel Architecture

Before an autonomous robot takes an action in the physical world, it should be able to dream.

Think about how humans navigate everyday life:

  • Catching a ball: When someone tosses a baseball to you, you don’t wait for it to touch your glove before reacting. Your brain automatically forecasts where the ball will be half a second into the future, and you position your hand where the ball is going to land.
  • Walking in the dark: When you turn off the lights in your bedroom at night, you can still walk to the doorway without crashing into the bed. You can’t see anything, but your brain maintains a working mental map of physical space.
  • Playing chess: A good player doesn’t move pieces randomly. They imagine: “If I move my knight here, my opponent will probably move their bishop there, and three moves later I’ll have the advantage.” They test hypothetical futures in their head before touching the board.

In Artificial Intelligence, this mental simulator is called a World Model.

1. What is a World Model?

In simple terms, a World Model is an AI agent’s internal imagination engine. It follows three basic steps:

  1. Compress: The AI takes sensor readings (camera frames, joint angles, or motor speeds) and condenses them into a compact thought representation called a latent state.
  2. Imagine: The AI tests hypothetical actions inside its mind: “If I push the motor with 2 Newtons of torque to the left, what will happen 1 second from now?”
  3. Plan and Act: The AI simulates dozens or hundreds of possible futures in milliseconds, picks the action that leads to the best outcome, and executes that move in reality.

Instead of trial-and-error in the real world, where making mistakes means crashing a drone or breaking a robot arm, the AI makes all its mistakes safely inside its imagination.

2. The Big Problem: Why Normal World Models Fail

Modern deep learning world models (such as Dreamer or MuZero) are remarkably good at video games like Atari or Minecraft. But when you try to use them on physical machines, they run into a serious issue: they hallucinate physics.

Standard neural networks treat the future as a pure pattern-matching game. They have no built-in understanding of gravity, momentum, or energy.

This leads to four major failure modes:

  • Compounding Errors: A 2% error on step 1 becomes a 20% error on step 5, and complete nonsense by step 15.
  • Teleportation and Infinite Speed: The AI imagines that an object suddenly gains huge speed with zero applied force, or unphysically stops in mid-air.
  • Vanishing Gravity: The model forgets that gravity constantly pulls downward, dreaming that an arm can float weightlessly.
  • Robot Hardware Crashes: If an AI dreams that gravity doesn’t exist, the action it chooses will fail completely when sent to real robot motors.

When you ask an unconstrained neural network to predict the future, you are essentially asking it to reinvent Isaac Newton’s laws from scratch every single time.

3. The Solution: Simple Physics Guardrails

Rather than letting the neural network guess whatever it wants, PI-WorldModel builds the fundamental rules of motion directly into the AI’s internal layers.

We do this with three simple design choices:

Guardrail 1: Separate Position from Speed

In normal world models, the internal state is one big unstructured array of numbers.

In PI-WorldModel, we explicitly split the AI’s internal state into two clear parts:

  • Where the object is (generalized position)
  • How fast it is moving (generalized velocity)

Guardrail 2: The Neural Network Only Predicts Acceleration

This is the key architectural shift. We do not let the neural network predict future positions or future speeds.

Instead, the neural network is only allowed to answer one specific question:

“Given the current state and the applied motor torque, what is the resulting acceleration?”

Once the model predicts that acceleration (aa), we do not use AI to figure out the next state. We use basic kinematics that everyone learns in school:

Next Speed=Current Speed+a⋅Δt\text{Next Speed} = \text{Current Speed} + a \cdot \Delta t

Next Position=Current Position+Current Speed⋅Δt+12a⋅Δt2\text{Next Position} = \text{Current Position} + \text{Current Speed} \cdot \Delta t + \frac{1}{2} a \cdot \Delta t^2

Because these formulas are hardcoded into the network, it is mathematically impossible for the model to hallucinate position without speed, or change speed without acceleration.

Guardrail 3: Work-Energy Balance

In the real world, you cannot get free energy. If a motor does work, kinetic energy increases. If friction slows things down, kinetic energy decreases.

During training, we penalize the model whenever the change in its internal kinetic energy doesn’t match the work done by the motor minus friction. This prevents the model from dreaming up runaway motion or free acceleration.

4. The Architecture: Tiny, Focused, and Fast

Because the model doesn’t need to learn basic calculus from scratch, it doesn’t need millions of parameters.

The entire PI-WorldModel has just 14,025 parameters, taking up only ~55 KB of memory.

Sensor Input [cos(theta), sin(theta), velocity]
                      │
                      ▼
            [ State Encoder ]          (4,676 params)
                      │
                      ▼
          Latent State: [Position, Speed]
                      │
                      ├──────────────────────────┐
                      ▼                          ▼
          [ Neural Acceleration Net ]     [ Applied Action ]
          (Predicts acceleration only)       (Motor Torque)
          (4,674 params)
                      │
                      ▼
          [ Kinematic Integration ]       (Exact Math: 0 params)
          Calculates new speed & position
                      │
                      ▼
          [ State Decoder ]               (4,675 params)
                      │
                      ▼
        Decoded Prediction: [Next Angle, Next Speed]

Teaching the Model About the Real World

To train the model effectively, we collected data using five distinct excitation patterns:

  1. Frequency Sweeps (Chirp): Gradually speeding up oscillations to teach the model about natural resonant frequencies.
  2. Harmonic Waves (Multisine): Combining multiple wave patterns to explore complex rhythms.
  3. Bang-Bang Control: Switching quickly between full-forward and full-backward torque to teach it about extreme motor limits.
  4. Sharp Shocks (Impulse): Quick bursts of power to observe sudden kicks.
  5. Smooth Random Walks: Natural, wandering motions that explore everyday operating states.

This ensures the AI learns how physical objects behave under gentle pushes, hard stops, and everything in between.

5. Planning in Pure Imagination: Latent MPC

Once the world model is trained, the agent doesn’t need to touch the real world to figure out what to do. It runs Model Predictive Control (MPC) inside its imagination.

Here is what happens during every single control step:

  1. Read Sensors: The robot reads its current angle and speed.
  2. Encode: Converts those readings into internal position and velocity coordinates.
  3. Dream 128 Futures: The agent creates 128 hypothetical action sequences and projects each one forward 12 steps in imagination.
  4. Score the Dreams: It evaluates each imagined future:
    • Is the pendulum upright?
    • Is it moving smoothly or shaking violently?
    • Did it use too much battery/motor torque?
  5. Execute the Best Move: It picks the first action from the winning trajectory, sends that torque command to the physical motor, and repeats the cycle for the next step.

We tested three different planning strategies:

  • MPPI (Path Integral): Blends promising action sequences using probability weights. It is exceptionally smooth and runs in 3.4 ms on CPU.
  • CEM (Cross-Entropy Method): Iteratively filters out the best candidate actions over several rounds. Highly accurate, running in 10.2 ms on CPU.
  • Random Shooting: Randomly samples actions and picks the lowest-cost path. Simple baseline running in 2.8 ms on CPU.

6. Microsecond Edge Execution: Pure C11

A major bottleneck in robotics is that research code is often written in Python and PyTorch. Python has background garbage collection pauses and heavy dependencies, making it ill-suited for real-time microcontrollers operating at kilohertz frequencies.

To solve this, I ported the entire neural network and MPC planner into pure C11 (c_runtime/world_model.c):

  • Zero Memory Allocation (malloc): The C code never allocates memory dynamically. Everything runs in static, pre-allocated memory or directly on the stack. No memory leaks, no fragmentation.
  • Total Size: ~55 KB: The compiled model weights and scratch buffers fit easily into the internal RAM of an STM32, ESP32, or ARM Cortex-M microcontroller.
  • Execution Speed: A single simulation step takes just 2.94 microseconds on a standard CPU core: roughly 340,000 steps per second.
  • Full MPC Loop: Evaluating 128 candidate futures over 10 time steps takes only 8.00 ms in pure C.

Here is a glimpse of how the kinematic update looks in C:

/* Exact Kinematic Step in C - zero dynamic memory */
float v_next = v + a_pred * dt;
float s_next = s + (v * dt) + (0.5f * a_pred * (dt * dt));

z_next[0] = s_next;
z_next[1] = v_next;

7. Performance & Comparison

How does PI-WorldModel compare against standard methods?

Runtime Performance

MetricPyTorch (CPU)Pure C11 EngineAdvantage
Model Parameters14,025 (~55 KB)14,025 (~55 KB)Identical
Single Step Latency230 microseconds2.94 microseconds78x faster
Single-Core Speed4,350 steps/sec340,100 steps/sec78x higher
Memory Allocation (malloc)Standard heapZero (100% static stack)Predictable timing
Kinematic Error0.00 (Exact)0.00 (Exact)No drift
Energy Drift (500 steps, unforced)< 0.0001%< 0.0001%Conserved

PI-WorldModel vs. The Unconstrained Baseline

To see what physical guardrails actually achieve, we compared PI-WorldModel against a standard black-box neural network with the exact same number of parameters:

  • The Black-Box Model: After 5 to 10 steps of imagination, it begins to drift. By step 15, it imagines the pendulum spinning at impossible speeds with almost no torque applied. When used for control, it repeatedly fails to balance because its imagined futures don’t match reality.
  • PI-WorldModel: Stays bounded and stable even across long horizons. The pendulum smoothly swings up from hanging down (θ=π\theta = \pi) to perfectly upright (θ=0\theta = 0) in under 1.5 seconds.

8. Where This Fits in the World Model Landscape

World models have evolved through several major milestones:

  • World Models (Ha & Schmidhuber, 2018): First popularized learning policies inside hallucinated dream states. Focused on video game pixels.
  • DreamerV1-V3 (Hafner et al., 2019-2023): Mastered games like Minecraft using recurrent state-space models. Powerful, but computationally heavy and unconstrained by physics.
  • MuZero (DeepMind, 2020): Planned moves in Go and chess using abstract latent search, without grounding states in real-world dimensions.
  • PI-WorldModel (This Work): Designed specifically for physical devices. Tiny (14k parameters), microsecond-fast, and mathematically grounded so it never dreams impossible motion.

When building AI for physical systems (robots, drones, factory machines, or prosthetic limbs), the goal isn’t to build a massive model that guesses everything. The goal is to build a reliable model that respects reality.

9. Quick Start: Try It Yourself

The entire project is open-source and runs directly on a standard laptop CPU without needing a GPU.

Clone and Install

git clone https://github.com/JashT14/PI-WorldModel.git
cd PI-WorldModel
pip install -r requirements.txt

Run the Complete Pipeline

Collects data, trains the model, tests multi-step imagination, runs closed-loop MPC control, and saves an animated visualization:

python main.py

Run the Interactive Graphical Simulator

python scripts/simulate_live_gui.py --interactive

Run the Standalone C Engine

# On Windows
./c_runtime/embedded_world_model.exe

# On Linux or macOS
gcc -O3 c_runtime/main_embedded.c c_runtime/world_model.c -o embedded_wm -lm
./embedded_wm

10. Summary

  1. Unconstrained models hallucinate: Pure pattern-matching neural networks invent impossible motion when forecasting physical systems.
  2. Split state into position and speed: Grounding internal coordinates prevents coordinate drift.
  3. Predict acceleration, calculate the rest: Let the neural network predict force and acceleration, and use standard calculus for the rest.
  4. Tiny footprint: At 14,025 parameters (~55 KB), you don’t need a cloud cluster to run intelligent physics simulations.
  5. Real-time edge control: Running in 2.94 microseconds with zero memory allocations makes world models practical for real microcontrollers and robots.

Sources & References