Protocol
Round 3 pilot cell. A dark 16:9 operations dashboard with fixed, machine-checkable content. Deterministic targets (graded by a separate grader, never the generator): exact OCR strings must all be present and correctly spelled: 'ARTBENCH OPS', 'STATUS: NOMINAL', 'UPTIME', '99.98%', 'LATENCY', '142 ms', 'ERRORS', '3', 'THROUGHPUT', '1,204 rps', 'REQUESTS / MIN'. Metric-card count must equal exactly 4 in a single row. Layout must resolve into three stacked horizontal zones top-to-bottom: title bar, one card row, one chart panel. Prompt bytes are frozen; the prompt_sha256 is recorded in docs/ROUND3_PLAN.md.
Frozen prompt template
CONTROL PROMPT R3-UI-DASHBOARD-001. A 16:9 dark operations dashboard user interface, flat modern design, near-black charcoal background, one full screen, no browser chrome, no mouse cursor. A top bar spanning the full width contains the left-aligned title text 'ARTBENCH OPS' and the right-aligned text 'STATUS: NOMINAL'. Below the top bar is a single horizontal row of exactly four equal metric cards, left to right, each a rounded dark rectangle showing a small uppercase label above one large numeric value: card 1 label 'UPTIME' value '99.98%'; card 2 label 'LATENCY' value '142 ms'; card 3 label 'ERRORS' value '3'; card 4 label 'THROUGHPUT' value '1,204 rps'. Below the four cards is one wide rectangular line-chart panel with the title 'REQUESTS / MIN' and a single cyan trend line over a faint grid. Colors are limited to charcoal, muted cyan, and off-white. All text is crisp, horizontal, correctly spelled, and sans-serif. No photographs, no 3D, no heavy drop shadows, no watermark, no extra panels, no decorative icons, no additional text beyond the labels and values specified.Review rubric
- Prompt Adherence
- Typography Fidelity
- Ocr Exact Match
- Count Accuracy
- Layout Fidelity
- Artifact Severity
Notes
Screening pilot executed 2026-07-26 with two unseeded repetitions for each available OpenAI endpoint policy. Dashboard OCR/card/layout deterministic extractors were deferred; published rubric rows are explicitly separate 0-5 visual-review scores. All four returned dashboards include extra axis text, recorded as a failed additional-text-constraint measurement. xAI cells are incomplete because the authenticated image route returned HTTP 403 and the stop gate prevented further retries. No global ranking may be claimed at N=2 reps per completed cell.