RoboJudge

Multimodal Language Models as Judges for Embodied Video Generation

Siyuan Ma*†1,2, Yu Qi*3, Cheng Ye*4, Ruotian Peng2,5, Xinyi Xu3, Zitiantao Lin3, Zedong Cai6, Yijian Huang6, Weiyang Liu1, Siyuan Wang✉1, Yandong Wen✉2
*Equal contribution.   †Project Lead.   ✉Corresponding author.
1The Chinese University of Hong Kong 2Westlake University 3Northeastern University 4University of California, San Diego 5Zhejiang University 6Independent

Human-aligned evaluation of physical adherence and instruction alignment in generated embodied videos.

RoboJudge benchmark, evaluation rubric, statistics, training data, and model overview
RoboJudge evaluates embodied video generation with human-aligned Physical Adherence and Instruction Alignment judgments.

Demo

Representative RoboJudge-Bench examples spanning complementary PA/IA outcomes. Videos load only when played.

Physically strong, partially aligned

Instruction: Sort the cotton swab containers and white packages from the two white rectangular containers into the blue tray on the table in the warehouse with the right gripper.

PA 5/5IA 3/5

Instruction aligned

Instruction: Crack the brown egg with your right hand over the black frying pan.

PA 3/5IA 5/5

Both axes fail

Instruction: Pick the dark glass from the wooden table and place it into the gray woven basket with the right gripper.

PA 1/5IA 1/5

Abstract

Video generation models are promising foundations for embodied world models, but only if their generated rollouts are physically valid and faithful to task instructions. Multimodal large language models (MLLMs) are increasingly used to judge these properties, yet their agreement with human judgments on embodied videos has not been systematically evaluated. To address this gap, we introduce RoboJudge-Bench, which measures Physical Adherence and Instruction Alignment across 800 human-annotated videos from 8 source corpora and 10 video generators, spanning grippers, human hands, and dexterous hands. Existing general-purpose MLLMs show limited agreement with human judgments, while specialized video judges transfer poorly to embodied scenarios. We therefore construct RoboJudge-Train, comprising 12K videos with 25K human-verified scores and rationales, and train RoboJudge-9B with supervised learning followed by reinforcement learning. RoboJudge-9B reaches an overall Pearson correlation of 0.719 with human ratings, an absolute improvement of 0.089 over the strongest proprietary baseline in our evaluation, and also improves on VideoPhy2 and Physion-Eval.

RoboJudge-Bench

Two complementary axes separate whether a rollout is physically believable from whether it performs the requested task.

PA · 1–5

Physical Adherence

Does the generated rollout obey the physical and temporal constraints of the scene?

  • Agent consistency
  • Scene consistency
  • Interaction realism
IA · 1–5

Instruction Alignment

Does the rollout preserve the intended agent, manipulate the correct object, and achieve the goal?

  • Agent match
  • Object correctness
  • Goal completion
800human-annotated test videos
8embodied source corpora
10test video generators
29MLLM judges evaluated

RoboJudge-Train & RoboJudge-9B

RoboJudge-Train contains 12K embodied videos annotated with 25K human scores and rationales for PA and IA. The released training annotations contain 12,351 PA records and 11,520 IA records. RoboJudge-9B is the resulting embodied-specialized multimodal judge; model weights and video assets are hosted on Hugging Face.
Stage 1

Supervised Fine-Tuning

Full-parameter language-model training with the vision tower and multimodal projector frozen, using the released PA and IA rationales.

Stage 2

GRPO Refinement

Multimodal GRPO with a schema-aware ordinal reward for exact and adjacent PA/IA score agreement.

Judge Leaderboard

Agreement with human ratings on RoboJudge-Bench

Top 10 model judges ranked by pooled Overall Pearson correlation on the same 800-video benchmark.

#1RoboJudge-9B0.719Overall Pearson r · Ours
#2Gemini-3.7-Flash0.630Overall Pearson r · Closed-source
#3GPT-5.20.626Overall Pearson r · Closed-source
RankJudgePA rIA rOverall r
1RoboJudge-9B OURS0.6450.7750.719
2Gemini-3.7-Flash CLOSED0.6110.7090.630
3GPT-5.2 CLOSED0.4990.7220.626
4Seed-2.1-Lite CLOSED0.5900.6940.619
5GPT-5.5 CLOSED0.4810.7310.607
6Claude-Opus-4.8 CLOSED0.5050.6550.587
7Qwen3.5-27B OPEN0.4310.6760.573
8Gemini-3.5-Flash CLOSED0.5940.5350.562
9Qwen3.7-Plus CLOSED0.5480.5460.538
10Qwen3.7-Max CLOSED0.5530.5390.537
11Gemini-3.1-Pro CLOSED0.4510.5840.529
12PhyJudge-9B JUDGE0.4170.5850.489
13Qwen3.5-9B OPEN0.3790.4610.401
14Cosmos3-Nano-Reasoner EMBODIED0.2560.5240.369
15Cosmos-Reason2-32B EMBODIED0.2610.4840.369
16GLM-4.6V-Flash OPEN0.2530.4400.356
17MiMo-Embodied-7B EMBODIED0.2090.4470.304
18Qwen3-VL-8B OPEN0.2920.3410.295
19Claude-Haiku-4.5 CLOSED0.1790.3570.278
20InternVL3.5-38B OPEN0.2040.3750.271
21MiMo-VL-7B OPEN0.0690.3620.237
22Cosmos-Reason2-2B EMBODIED0.0850.3090.205
23VideoScore2-8B JUDGE0.1950.1910.202
24RynnBrain-1.1-9B EMBODIED0.0510.1960.129
25RoboBrain2.5-8B EMBODIED0.0420.1970.122
26VideoPhy2-7B JUDGE0.1520.1470.103
27Qwen3.5-4B OPEN0.1390.2350.101
28VideoCon-7B JUDGE0.0950.0550.077
29Orca-4B EMBODIED0.1080.0960.059
—WBench-30B-A3B JUDGE0.412——

+0.089 overall Pearson r over Gemini-3.7-Flash.