MaLiang-Harness generates images and videos by having a multimodal LLM write and iteratively revise executable drawing code, keeping a versioned program state so the model can inspect each rendered output, locate the offending code, and edit it rather than regenerate from scratch.
Most image and video generation today uses diffusion or flow-matching: the model paints pixels directly from a prompt, and if the result is wrong you re-prompt and hope. You cannot point at “the vase on the left” and ask the model to move it 20 pixels right while keeping everything else identical, because there is no explicit representation of “the vase.”
An alternative is to have an Multimodal Large Language Model write code (Canvas, SVG, Three.js, or a path tracer) that draws the scene, then render it. This gives you controllability and editability for free. But the authors point out a problem they name the Program-to-Visual (P2V) gap: the code can run without errors and still produce the wrong picture. The object is in the wrong place, the animation fires at the wrong moment, the composition is off. Execution success is not the same as visual-requirement success, so you need a loop where the model looks at the rendered output and revises the code.
Prior work like VISPROG and BlenderAlchemy already showed programs can mediate visual tasks. MaLiang-Harness is positioned as the scaffolding that makes this loop stateful across many revisions and multiple rendering backends.
The contribution is three coordinated bookkeeping mechanisms, not a new model. Think of it as git-for-generated-artwork, with mandatory review gates.
•
Persistent Executable Generation (PEG) state. At every revision the harness stores the full program, any image assets, the output spec (dimensions, seed, timing), and the task context (prompt, user requirements, current plan). Each edit produces a new numbered revision; old revisions are kept, not overwritten. The model always has a stable, named object to point at.
•
Traceable Generation Process (TGP). Every operation the model performs (edit, render, inspect) is logged with its input, output, source revision, and resulting revision. Rendered observations are tied to the exact revision that produced them. When a picture looks wrong, you can walk back through the operation log and find which code change broke it.
•
Revision-aware Editing and Verification (REV). The model maintains a review for each visual requirement: which revision was checked, what evidence (full renders, crops, or ordered video frames at 3+ timestamps) was used, and a pass/fail/uncertain verdict. Crucially, a review only counts if it was done on the current revision. Any commit, even a plan-only change, invalidates prior reviews and requires re-verification before the task can be marked complete.
In plain terms: the readiness check for “done” is export validity AND every mandatory requirement passing a review bound to the current revision. The model can also restore a prior revision as a new revision (preserving current requirements), which supports “undo” without losing the task context.
state = PEGState(program=None, assets={}, context=prompt_and_spec, rev=0)
while not ready(state) and budget_remaining():
plan = mllm.plan(state)
edit = mllm.generate_edit(state, plan) # code or asset change
state = commit(state, edit) # new revision, old kept
image = render(state.program, backend)
log_observation(image, state.rev)
for req in state.context.requirements:
review[req] = mllm.verify(image, req, state.rev) # must match current rev
return export(state) if ready(state) else draft(state)
The evaluation runs 11 closed-source MLLMs through the harness on MaLiang-IBench (50 text-to-image prompts) and 4 on MaLiang-VBench (13 text-to-video prompts). A separate model, GPT-6-Sol, acts as judge on a 1-5 scale over alignment, aesthetics, composition, and (for video) motion coherence. The denominator for pass rates is the full task set, so failures count against you.
•
Execution success is not visual success. GPT-5.6-Luna and GPT-5.6-Terra both successfully generate 46/50 images, but only 22 and 24 meet all three quality thresholds. The authors treat this gap as direct evidence of the P2V problem they named.
•
The strongest model clears both bars. gpt-6-astra hits 100% generation success on both benchmarks, with 96.0% of images and 76.9% of videos meeting every quality threshold. For comparison, GPT-5.6-Sol hits 86.0% and 38.5% on the same “all criteria” measure.
•
Weaker models mostly fail to even produce output. DeepSeek and Kimi variants complete 12-40% of image tasks and 0-23% of video tasks. Their mean quality scores on the few successes are not comparable to GPT numbers because the denominators are tiny (6-20 images).
•
Motion is the hardest video criterion. All 13 Astra videos clear the aesthetics threshold but only 10 clear motion coherence.
•
General capability doesn’t fully predict visual programming skill. Comparing against public Artificial Analysis Intelligence Index scores, Spearman rank correlation is ρ=0.65. But GPT-5.6-Luna and GPT-6-Luna share an index score of 37, yet their “all criteria” pass rates are 44% vs 88%. Kimi K3 scores 44 on the index but passes all visual criteria on only 18% of tasks.
•
Refinement can regress. Qualitative examples show local edits introducing new artifacts (overlap glitches on a biscuit after a geometry change), and a brushwork-refinement experiment made some regions less detailed, not more. PEG+TGP+REV let the model notice and sometimes fix these regressions, but don’t guarantee it.
Important caveat on the cost numbers: the paper itself warns that per-task time, call, and token aggregations use different denominators across models (some use successful tasks, some use the full set, some use medians), so direct efficiency comparisons across rows are not apples-to-apples.
•
If you’re building a code-generating visual agent, copy the revision discipline. The core practical insight is binding every verification review to a specific revision ID and invalidating reviews on commit. Teams building agentic image or UI generation often skip this, which lets the model declare success based on stale observations. Worth adopting even without the rest of the framework.
•
Pick your judge carefully if you reuse this evaluation style. GPT-6-Sol is both evaluated and used as judge in the same paper. If you run a similar benchmark, consider a judge from a different family to reduce the “same model grading its cousins” concern. The paper notes the review was not formally blinded.
•
Don’t use general-capability leaderboard scores as a proxy for visual-program generation. The 44pp gap between two models with the same AA index is the strongest practical lesson: if your product needs code-to-image, benchmark it directly, don’t trust Arena/Index rankings.
•
Worth testing on your own stack: whether swapping the rendering backend (Canvas vs path tracer vs SVG) changes which failures the model can recover from. The paper shows a path-tracing example where richer rendering helps on materials but introduces new local defects. Controlled comparisons are not reported.
•
The model names used (GPT-6-Astra, GPT-5.6-Sol, Kimi-K3, DeepSeek-V4-Pro, etc.) and the AA Index snapshot are dated September 27, 2026. Treat specific numeric comparisons as a snapshot of that generation, not a durable ranking.
•
The judge is an MLLM, not human raters. The authors explicitly say “MLLM verdicts remain self-assessments rather than independent measures of perceptual quality.” The one human-reviewed slice (23 videos) was not blinded.
•
Benchmarks are small: 50 image prompts and 13 video prompts. Per-dimension counts can swing on a handful of tasks.
•
The harness itself is scaffolding; it does not improve a weak model. DeepSeek and Kimi variants fail mostly because they cannot produce workable code within the token budget, not because they lacked state management.
•
Photorealism remains out of reach via Canvas drawing routines alone. The brushwork-refinement experiment shows that asking for “finer strokes” can degrade a stylized illustration without approaching photorealism.
•
Cost aggregation rules differ across rows (per-task vs per-success vs per-qualified denominators, medians vs means), so time/token comparisons across models should be read as rough, not precise.