The user never gave the value. Does the model write one?
Pick a scenario and a model. Each card is that model's real output under one condition, with the field the user never filled marked: red where the model wrote a value, blue where it left the field empty. Verdicts are the paper's strict scoring rule, recomputed here from the released outputs.
All scenarios for this model
Columns: prose, required, nullable, required + license, nullable + license. Click a square to open it.
Permission, not nullability
Click a cell to highlight it in the per-model view. Rows: whether the prompt adds one sentence of permission (the license). Columns: whether the field is required or nullable in the schema.
Languages, real dialogues, native tool calls, other groups' benchmarks
The same comparison, run again where the paper tests generality. Hover a bar for its count.
It is the content of the sentence, not its position or a hint from examples
Where the decision to abstain is made
One per level of access
Every figure in the paper
Rendered from the submitted figure files. Click one to enlarge it with its caption.
Every table in the paper
Everything needed to check the numbers
Rebuild every table and figure (CPU, about a minute)
pip install numpy scipy matplotlib pillow python-dotenv openai python3 scripts/make_tables_figures.py python3 scripts/make_schema_general.py python3 scripts/fig1_arch.py schema python3 scripts/fig_ladders.py schema
Run from the top folder of the archive. The regenerated paper/tables/*.tex match the
shipped copies byte for byte. Re-running an experiment needs API keys (read from environment variables) or a GPU;
see the README inside the archive.