Physical Intelligence π0.7 Explained: A Steerable VLA for Unseen Robot Tasks

Physical Intelligence's official π0.7 article describes a vision-language-action model whose context can steer the manner of execution, not only name the goal. Language coaching, desired quality or speed metadata, and generated subgoal images can disambiguate how the same task should unfold.

The company reports unseen-task composition, cross-embodiment transfer and specialist-level results on named experiments. Those results should remain attached to the robots and tasks tested. For the underlying model class, read our VLA versus VLM guide; this page explains what is new about π0.7's steering.

Steerable means specifying the method as well as the goal

A standard instruction such as “clean the kitchen” leaves many strategies unresolved. π0.7 can receive intermediate language, strategy metadata and visual subgoals that describe desired progress, letting the policy distinguish among behaviors represented in a mixed dataset.

This is not the same as changing a motor parameter with ordinary prompting. The model was trained with those context types, sometimes dropped out, so the signals can shape a learned action policy at inference.

Context signalWhat it specifiesExample use
Task instructionOverall goalClean the counter
Subtask languageNext semantic stepOpen the drawer
Episode metadataQuality, speed or strategyPrefer a fast successful style
Subgoal imageDesired near-future visual stateShow the intended object layout

The model combines several kinds of experience

The paper describes training on robot demonstrations, autonomous episodes including failures, egocentric human video and multimodal web data. Rich context labels help the model avoid averaging incompatible strategies into an indistinct behavior.

Its architecture builds on the π*0.6 and Recap line and MEM, with a visual-language backbone, video-history memory and an action expert. Generated subgoal images come from a lightweight world model and complement language when spatial details are hard to express.

Language coaching can become an autonomous high-level policy

In the π0.7 paper's air-fryer experiment, a short zero-shot instruction produced only partial progress. Step-by-step verbal coaching improved execution, and the recorded coaching was then used to fine-tune a high-level policy that generated subtasks autonomously.

That sequence matters: interactive coaching and later autonomous execution are separate stages. It would be misleading to show the coached trial and describe every step as independently planned by the base model.

Reported testObserved signalLimit to retain
Air fryerComposed a new appliance task with coachingInitial zero-shot attempt did not finish
Laundry on bimanual UR5eCross-embodiment transfer without task data on that robotOne named embodiment and task setting
Espresso, laundry and box tasksGeneralist matched or exceeded named specialistsCompany experiment, not every manipulation task
Open-ended commandsFine-grained language followingNot unrestricted safe natural-language control
Collage of Open X-Embodiment robot arms performing manipulation tasks
This image is a collage of Open X-Embodiment robot-task scenes, not a Physical Intelligence π0.7 experiment. It does not prove π0.7's steering behavior or unseen-task performance. Source: Open X-Embodiment robot dataset scenes. License: Project license and dataset-specific terms.

Unseen task does not mean outside all prior knowledge

The model may not have received a demonstration of the exact task on the evaluated robot, yet related objects, movements and semantic concepts can appear across robot, human and web data. Compositional generalization means recombining those ingredients, not acting without prior information.

When judging a claim, ask what was held out: the instruction, object, scene, embodiment or full task combination. Our robot VLA evaluation guide shows why these split definitions change the strength of the result.

The public material is research, not a turnkey API

The April 2026 article and paper describe the model and experiments and invite research or application collaboration. They do not publish a general self-service commercial API, a universal supported-robot list or guaranteed service levels.

A prospective integrator should request checkpoint or access terms, hardware requirements, control frequency, safety interfaces and robot-specific adaptation needs. Do not infer commercial availability merely because a detailed paper and videos are public.

Mobile decision card summarizing four key checks for Physical Intelligence π0.7 Explained: A Steerable VLA for Unseen Robot Tasks
A Physical AI Lab editorial card based on the article's cited official sources and comparison table. Source: Physical AI Lab. License: Owned original.

A local trial should test control, not just task completion

Create paired prompts that keep the goal fixed while changing speed, strategy or visual subgoal. Measure whether the physical trajectory changes in the requested direction without reducing success or violating workspace constraints.

Then test held-out objects and environments with repeated runs, logging recovery and intervention. The robot action-data guide can help define the trajectory record needed to distinguish genuine steering from selectively chosen clips.

Frequently asked questions

What makes π0.7 steerable?

Its trained context can include subtask language, strategy metadata and visual subgoals, allowing the prompt to specify aspects of how a task should be performed as well as the goal.

Did π0.7 complete every unseen task zero-shot?

No. The publication reports early signs of compositional generalization and includes an air-fryer attempt that improved with step-by-step language coaching.

Can developers call a public π0.7 API?

The official article and paper describe research and invite collaboration, but they do not announce a general self-service commercial API.

Official sources checked

Last checked: August 7, 2026