Training chatbots to use robots
If you're paying attention, you may have noticed that so-called LANGUAGE models are increasingly nailing visual tasks. Robotics, 3D design, image generation: all going the way of the transformer architecture.
Why is the chatbot absorbing these things too? This is because, to the computer, text is undoubtedly the universal interface. Everything that the computer understands is ultimately relayed to it via some kind of text. As such, tasks that we perceive as visual can be distilled into textual formats called structured state representations and trained upon to improve visual decision making.
This was always possible, but there have been obstacles for a language model to, say, predict motor values at every timestep for a robot. Here's a big one for robotics: these spatial tasks are slow and expensive when carried out by an LLM, but the task demands real-time reactivity. The LLM has to generate a decision for every step, or else follow some preset plan between refreshes that cannot be adjusted, all while burning tokens. This all amounts to overkill.
The System 1/System 2 paradigm of models breaks this limitation. Jev, the heralded System 1 model, can act as a scout, quickly assessing the changes in that structured state over time. Jev relays the changes to the slower, wiser System 2 model, which then rewrites what's left of the plan. Jev simply reads the state, then picks among a limited set of actions defined in the structured state. System 2 makes the initial plan and then, like a watchful guardian, ONLY steps in when Jev assesses from the structured state that there has been a failure in the step or that the task has changed considerably. My System 2 model was OpenAI's GPT-6 Sol. I tested Jev and Sol separately, and I haven't run the full loop yet.

With that, the structured state representation of the scene becomes very important for a language model to be able to work through these spatial tasks. The System 2 LLM encodes its decree for Jev within it, and then Jev reads it at every decision. With the LLM generating the high-level directives contained in the state, it can be done live to cover the System 2 role for a task at inference-time, or it can be done offline at scale to make lots and lots of training labels for some class of visual tasks that you can then improve upon.
The structured state
So what even is this structured state? What does it consist of? Everything? It's supposed to describe the entire scene in a compact way. I tried to understand what the best makeup for this structured state was, and how to create them at scale to train better models.
Let's start with an easy case. Say my robotic task is inside of a simulation engine. In my engine, I can get a report of every object's exact position at every step. This tracking is the bedrock of my structured state. In MuJoCo I read the positions straight from the engine, add a few mm of noise to mimic a real tracker, and store them relative to each other, like where the cube is relative to the gripper.
The state also needs a menu of all possible actions the robot can take at any given timestep. The model's whole job is to choose from it, so that menu constrains what the system can do at all. We define this menu ourselves. Mine were skills like "approach the cube", "close the gripper", "lift", "carry to the tray", "release" and "back off", and for drawing, "draw this line", "lift the pen" and "finish".
And then lastly, the goal label: what the task is and what counts as done. This is a very open-ended field, and yet we are going to need to find some kind of verifier that can test how complete a task is.
I then ran ablations across all of these fields to find which ones actually add meaningful functionality to the model. I ran a bunch of tasks with and without each field and checked whether its absence hurt the model, and at what degree of imprecision the field hurt the model when present.
What goes in the state
Through this, there were some interesting findings. Mostly that, as long as the goal was clearly stated, whenever I tried to add an additional label, it ended up being totally dominated in usefulness by the tracking data anyway, even accounting for tracking error and noise. When the goal was badly written, the extra labels did partly patch over it.
We could distinguish these experimental labels as:
- Meaning labels: simple yes/no facts computed from the positions: "gripper lined up with the cube", "cube grasped", "cube in the tray"
- Procedure labels: bookkeeping checks for the task in progress ("is my position estimate out of date?", "how many grab attempts since the last lift?")
These generally went nowhere. The meaning labels describing contact that did survive ablation would reveal a category I might call "fine contact" (i.e. "Is the pen touching the paper?") where a few millimeters of discrepancy can majorly impact the success of the task. The fallout from fine contact error was also lopsided, as a false "not touching" knocked the pen out of the hand every time, while a false "touching" did no harm. In a real task, these labels would require a precision that outstrips tracking error. If tracking error were higher than the precision needed, these "fine contact" interactions would underscore areas that would need actual additional haptic sensors to represent accurately.
This gradually whittled down the structured state into a leaner, more useful representation. But most of all, it identified the major defining term of our structured state.
The goal
The goal decided more than any other field, and no sensor can hand it to you. Somebody has to write down what the task is and when it's over.
A goal with a bad ending was worse than no goal at all. My first pick and place goal said "Pick up the red cube from the table and place it in the blue receptacle." It never said to let go and back away, which is what the checker wanted. With that goal and positions alone, Jev finished 17 of 50 runs, re-grabbing the cube until it ran out of decisions. Once the goal ended with "then move the gripper clear of the receptacle," positions alone got 50 of 50. With no goal at all, Jev got 49 of 50 (48 on a second set of seeds). A goal missing its ending landed between 17 and 22 of 50 four separate times.
This is where the extra labels earned their keep. Under the half-written goal, meaning labels lifted Jev to 28 and procedure fields got it to 47. Under the full goal, the same fields dropped it to 45, because Jev started looping.

No goal did well because the menu was narrow. With skills like approach, close, lift, carry, release and back off, there's really one task you can do, so the menu quietly carries the goal. Such a menu only works for that one task, though, and I set mine by hand.
Labeling goals from video
The obvious way to get goals for thousands of videos is to have a multimodal model watch each one and write the goal. Sol mostly wrote down what it saw. It left out endings: 1 of 50 pick and place goals mentioned the robot's final withdrawal, and 15 of 50 drawing goals mentioned lifting the pen. It copied failed attempts. After a button press where the first try missed, one goal read "Attempt to grasp the front red piece twice, then withdraw without moving it." And one unrelated red cube in the scene made 21 of 50 button goals describe a pick and place.

A wrong goal at scale gets copied into every label. The most reliable fix I found skipped the goal entirely: I built the ending into the robot's finish action, so "finish" always lifts the pen. Without it, 283 of 400 replayed drawings dragged a line at the end. With it, drawings ended cleanly whatever the goal said.
It isn't hopeless. After watching the robot draw one part of a painting, Sol named the part 12 of 12 times, and its goal got about halfway from "draw this image" to my own goal (44% of the ink inside the part, against 16% and 66%).
Splitting the goal up
The System 2 model's real job is turning the goal into a plan. Jev can follow a plan written in advance, like a list of strokes, and recovered from all 14 interrupted strokes I threw at it. Sol can write that kind of plan straight from a picture, though Jev hasn't run one of Sol's plans yet. In my drawing testbed, Sol looks at a painting and writes every stroke, and a robot hand from llm-robotics-playground draws them.
Models like π0.7 lean on sub-goals, so I tried breaking the job up. I wrote each painting's parts by hand and compared five ways: one shot, four rounds of the whole picture, four rounds of one part each, coarse to fine, and parts with a layout, where Sol first plans a box for each part. I rated 12 drawings per way, blind, from 1 to 5.
Breaking it up didn't help! One shot averaged 2.25, parts with a layout 2.17, rounds by part 1.67, rounds 1.08 and coarse to fine 1.00. One shot also used the fewest tokens, and a checklist of parts moved its rating by all of 0.07.
In this fresh run, the Starry Night rounds drawing looks about as good as one shot, while Mona Lisa keeps the old order. That's one unrated sample, so I can't say yet whether the current Sol closes the gap.
My guess is that rounds hurt because Sol only saw a picture of the page between rounds and struggled to place new strokes against old ones. Giving it the coordinates of its earlier strokes improved rounds on the automatic scores, but I haven't rated those drawings yet, so that's only a hint.
Parts drawn one at a time also lost track of each other. With no reference image, the Statue of Liberty drawn part by part came out as a pile. A sub-goal needs a "where" as well as a "what", which is what the layout added, though with 12 drawings each its edge over parts alone isn't established.
robolabel
All of this points back at the labeler. robolabel is my tool for turning demonstration video, from plain clips or LeRobot datasets, into training labels: the task's phases, each grasp attempt and whether it failed, and the goal. The new version works from video alone. Its goal is a list of end states, each marked required, incidental or unsure, and a failed grasp gets its own label as a failed attempt, which is aimed at the copied retries above. I haven't measured how accurate those goals are yet.
Where this goes
So what actually goes in a structured state? Less than I expected. Positions, even noisy ones. A menu of actions the robot can really take. A real contact sensor wherever a few millimeters decide the outcome. And a goal, stated completely, especially how it ends.
Next I want to run the whole loop, with Sol writing the plan and Jev running it, then train on goal labels like robolabel's. The tasks, the checks and a short history of all 19 rounds, mistakes included, are in the statebench repo.
Put together, that's the shape of an RL environment for spatial skills: a scene written as structured state, a checker that says whether the task got done (which doubles as the reward), and goal labels made at scale from video. That's how you could train a chatbot toward the spatial smarts newer models like GPT-6 Astra are showing off. I don't know how Astra itself was trained, but this is how I'd go about it.