ROBOT EVALUATION ARCHIVE

Watch the policy.
See every outcome.

A coding agent controlling robots through visual feedback. Browse recorded rollouts, from the first move to the final result.

Explore the collection GPT-6-Astra high reasoning
… tasks… rolloutsEvery success. Every failure.
Loading recorded rollouts…

Separate benchmarks, different tasks and control protocols. Success rates are reported independently.

THE COMPLETE RECORD

The rollout collection.

Download source CSV

Loading recorded episodes…

Loading tasks…

SuccessFailureTimeout

HOW TO READ THIS GALLERY

The task is visual.
The score is native.

These evaluations use GPT-6-Astra with high reasoning. Each episode starts with a fresh context and receives its native task instruction, RGB camera views and robot state. Camera calibration follows each benchmark’s protocol; RoboDojo exposes no calibration or depth.

Read the evaluation protocol
All scored episodes are included

The published selection includes successful and unsuccessful episodes. Source selection and infrastructure retries are documented in the evaluation protocol.

Success comes from the simulator

Results use each benchmark’s native success condition, not the agent’s written assessment. The policy receives no object ground-truth poses, reward or success labels, training demonstrations, or cross-episode memory. In interactive RoboDojo tasks, the native support arm’s demonstration remains visible as part of the task.

Playback and control time are different

The player describes each benchmark’s recorded cameras and playback rate. Thinking time is omitted from video playback.

Sampling and comparison

Sampling is reported separately for each published benchmark. Different tasks and protocols prevent direct comparison or pooling across benchmarks.