Robot-Use Agents: Why General-Purpose Models May Win in Robotics
Y Combinator · 29:49 · 2 days ago
General-purpose LLMs can already execute demoed robot manipulation from camera frames via tool-call harnesses, and the broader two-year robot timeline is framed by participants as a target rather than a measured benchmark .
Key Points
- Robot-specific VLA training is framed as the data bottleneck: RT-style VLAs used a web-pretrained language model to output end-effector poses, and web image/text pretraining improved them over a robot-only foundation model . Code-as-policy research then gave LLMs Python-style robot functions, allowing one-shot pick/lift/move task composition without additional robot data . In-context learning also improves quickly but saturates after roughly 20–40 examples, and it degrades once context exceeds the model’s effective window .
- Waddle Labs describes a harness that feeds camera images to a general model, which returns end-effector poses or tool calls rather than always writing full code . Direct model-in-the-loop control is too slow for repetitive motion, so repeated trajectories should be compiled into faster reusable skills . A VLM can detect objects and handle failures, and model latency was reported to improve about 2x per month, with real-time control claimed possible by year-end .
- RoboCurve’s role is evaluation rather than control: it benchmarks arbitrary robot embodiments—hands, grippers, arms, humanoids, quadrupeds—against LLMs, VLAs, and world-action models . The target’s capability bar is following natural-language instructions to do what a competent teenager can do with hands in unseen tasks and environments . Stated remaining work includes converting slow in-context learning into faster skills and weight updates, with sleep-like offline consolidation as one proposed mechanism .
Step by step how to
- Give a general model camera frames and the robot’s tool/API schema, rather than training a robot-only policy .
- Let the model output end-effector poses or tool calls, then convert those signals into joint commands .
- Compile repeated, deterministic motion segments into callable skills or code so the model only reasons at ambiguous steps .
- Use a VLM inside the harness for object detection and failure checks, allowing non-deterministic re-planning .
- After an episode, consolidate successful traces into faster skills or weight updates in a sleep-like offline pass .