Robot foundation models

HABILIS Brain 0

Our robot foundation model architecture,
demonstrated through sustained operation, precise manipulation, and adaptive planning.

Loading…

One foundation. Built for real robot work.

The foundation behind HABILIS Brain 0

Our robot foundation model architecture brings together visual reasoning, continuous action, and recovery during execution.

GC-VLA architecture: egocentric human video, UMI-style data, simulation and teleoperation feed GC-VLM and an action expert, with a unified 14D action space for single- and dual-arm robots.
01 · RepresentationGC-VLM

Learns representations of how a scene should change.

02 · ActionGC-VLA

Translates scene-change representations into continuous robot actions.

03 · RecoveryGCRF

Adds selective recovery during execution.

The applications below build on this architecture with task-specific learning, recovery during execution, and live planning.

Keeps working. Reliably.

Tote picking that keeps going across repeated clears. Operational experience is collected and used to retrain the policy for more varied object poses.

Loading…
HABILIS Brain 0Tote picking10× timelapseClear 01 / 10

What to watch

  • Sustains continuous operation Keeps clearing totes, cycle after cycle.
  • Repeats the full picking cycle Carries out picking and placement repeatedly as each tote is emptied.
  • Handles random object poses Picks objects in random positions and orientations after they are poured into the tote.
  • Learns from operational experience Improves through retraining on autonomous rollouts and human-intervention data.

Reliability loop

Learns from experience.

Start with task demonstrations, collect experience during operation, then retrain for more varied object poses.

01

Initial fine-tuning

Fine-tuned with just 3 hours of demonstrations

This initial task dataset covers a limited range of object poses. Rollout and intervention data have yet to be added.

Loading…

02

Rollout and intervention

Learning from rollouts and human corrections

Autonomous rollouts and human corrections are recorded as training data. The expanded dataset is then used to retrain the policy.

Loading…
01Autonomous rollout
02Human intervention
03Dataset update
04Retraining

03

Retraining results

Robust to random object poses

After retraining with rollout and intervention data, the robot handles varied object positions and orientations in this demonstration.

Loading…
HABILIS Brain 0Tote picking10× timelapseClear 01 / 10

Stays precise across long horizons.

Loading…
HABILIS Brain 0Glasslock packing
00:23

What to watch

  • Executes long-horizon tasks as one continuous sequence.
  • Maintains precision at each manipulation step.
  • Coordinates both arms throughout the task.
  • Recovers and continues when the task state changes.

01

Recovery during execution

Recovers from disruptions.

The policy responds to a changed task state, restores what is needed, and continues the current task.

Loading…
Recovery from object lossDetects the missing object, recovers the task state, and continues.
Loading…
Recovery from displacementDetects the shifted state, restores the task, and continues.

02

Human-guided reinforcement learning

Reinforces the right behavior. Fast.

In the session shown, online reinforcement learning with human guidance corrects the behavior in under 28 minutes.

Behavior correction sequence

Compare the same task before and after correction.

  1. Loading…
    Familiar starting positionHandles varied container positions across the brown tray, the intended starting area for this task.
  2. Loading…
    Before correctionStarting off the tray introduces a condition beyond the original task setup. The initial policy repeats an incorrect action.
  3. Loading…
    Human-guided RLAutonomous rollout episodes and human-taught trajectories are collected as training data. Reinforcement learning updates the actor policy in real time as this experience accumulates, refining how the robot selects its actions.
  4. Loading…
    After correctionThe robot learned the corrected behavior in under 28 minutes, demonstrating targeted correction of individual actions within a long-horizon task sequence.
Full learning session

Human-guided reinforcement learning, from initial attempts to corrected behavior.

Loading…
Reference01/04
HABILIS Brain 0 · Full learning session · 10× timelapse
00:00

Thinks. Plans. Reacts.

The robot breaks a goal into subtasks, detects meaningful changes, and updates its plan as the task changes.

Loading…
HABILIS Brain 0Block Disassembly & Sorting
00:00

What to watch

  • Breaks goals into subtasks Turns the overall task into actionable steps.
  • Reasons from the current scene Uses object and task state to decide what comes next.
  • Connects planning to action Guides the robot through each subtask during execution.
  • Responds to scene changes Adjusts its actions to fit the updated situation.
Loading…
HABILIS Brain 0Block Disassembly & SortingReal-time perturbation

Context-aware behavior

Recognizes change.
Reacts immediately.

When object colors change, the model recognizes the updated scene and adjusts its actions to suit the current situation.