Herman was retired on September 10, 2026. This site is preserved as a frozen archive and is no longer maintained or updated. Read the retrospective →

Experiment · active

Round 3 · Dark operations dashboard fidelity

At a fixed dark-dashboard brief, how faithfully do hosted endpoints reproduce exact labels, a four-card count, and a three-zone layout?

Runs
86
Variable
model endpoint or quality tier
Prompt
r3-ui-dashboard-fidelity-prompt v1.0

Protocol

Round 3 pilot cell. A dark 16:9 operations dashboard with fixed, machine-checkable content. Deterministic targets (graded by a separate grader, never the generator): exact OCR strings must all be present and correctly spelled: 'ARTBENCH OPS', 'STATUS: NOMINAL', 'UPTIME', '99.98%', 'LATENCY', '142 ms', 'ERRORS', '3', 'THROUGHPUT', '1,204 rps', 'REQUESTS / MIN'. Metric-card count must equal exactly 4 in a single row. Layout must resolve into three stacked horizontal zones top-to-bottom: title bar, one card row, one chart panel. Prompt bytes are frozen; the prompt_sha256 is recorded in docs/ROUND3_PLAN.md.

Frozen prompt template

CONTROL PROMPT R3-UI-DASHBOARD-001. A 16:9 dark operations dashboard user interface, flat modern design, near-black charcoal background, one full screen, no browser chrome, no mouse cursor. A top bar spanning the full width contains the left-aligned title text 'ARTBENCH OPS' and the right-aligned text 'STATUS: NOMINAL'. Below the top bar is a single horizontal row of exactly four equal metric cards, left to right, each a rounded dark rectangle showing a small uppercase label above one large numeric value: card 1 label 'UPTIME' value '99.98%'; card 2 label 'LATENCY' value '142 ms'; card 3 label 'ERRORS' value '3'; card 4 label 'THROUGHPUT' value '1,204 rps'. Below the four cards is one wide rectangular line-chart panel with the title 'REQUESTS / MIN' and a single cyan trend line over a faint grid. Colors are limited to charcoal, muted cyan, and off-white. All text is crisp, horizontal, correctly spelled, and sans-serif. No photographs, no 3D, no heavy drop shadows, no watermark, no extra panels, no decorative icons, no additional text beyond the labels and values specified.

Review rubric

  • Prompt Adherence
  • Typography Fidelity
  • Ocr Exact Match
  • Count Accuracy
  • Layout Fidelity
  • Artifact Severity

Notes

Screening pilot executed 2026-07-26 with two unseeded repetitions for each available OpenAI endpoint policy. Dashboard OCR/card/layout deterministic extractors were deferred; published rubric rows are explicitly separate 0-5 visual-review scores. All four returned dashboards include extra axis text, recorded as a failed additional-text-constraint measurement. xAI cells are incomplete because the authenticated image route returned HTTP 403 and the stop gate prevented further retries. No global ranking may be claimed at N=2 reps per completed cell.

Evidence

Recorded runs

errorGrok Imagine

R3 Ui Dashboard Fidelity Grok Imagine Trial 01 Failure

Artifactless provider-route failure. HTTPError: HTTP Error 403: Forbidden. No visual model output exists to score.

Size
not returned
Seed
not exposed
Method
first party tool
View experiment →
Generated specimen r3-ui-dashboard-fidelity-openai-gpt-image-2-high-01 r3-ui-dashboard-fidelity-openai-gpt-image-2-high-01
successGPT Image 2 High

R3 Ui Dashboard Fidelity Openai Gpt Image 2 High 01

Screening grade: partial. High-01 dashboard hits all required OCR strings, four-card order, and three stacked zones on charcoal/cyan/off-white; only material miss is unspecified chart axis tick/time text. Rubric: artifact-severity=4/5, count-accuracy=5/5, layout-fidelity=5/5, ocr-exact-match=5/5, prompt-adherence=4/5, …

Size
1672×941
Seed
not exposed
Method
official api
View experiment →
Generated specimen r3-ui-dashboard-fidelity-openai-gpt-image-2-high-02 r3-ui-dashboard-fidelity-openai-gpt-image-2-high-02
successGPT Image 2 High

R3 Ui Dashboard Fidelity Openai Gpt Image 2 High 02

Screening grade: partial. High-02 matches title/status, exact four-card metrics, chart title, and three-zone structure; residual issue is extra chart axis numerals not listed in the frozen prompt text set. Rubric: artifact-severity=4/5, count-accuracy=5/5, layout-fidelity=5/5, ocr-exact-match=5/5, prompt-adherence=4/5, …

Size
1672×941
Seed
not exposed
Method
official api
View experiment →
Generated specimen r3-ui-dashboard-fidelity-openai-gpt-image-2-medium-01 r3-ui-dashboard-fidelity-openai-gpt-image-2-medium-01
successGPT Image 2 Medium

R3 Ui Dashboard Fidelity Openai Gpt Image 2 Medium 01

Screening grade: partial. Medium-01 reproduces all machine-checkable dashboard strings and the four-card three-zone layout cleanly; only clear constraint slip is residual y-axis tick copy on the chart. Rubric: artifact-severity=4/5, count-accuracy=5/5, layout-fidelity=5/5, ocr-exact-match=5/5, prompt-adherence=4/5, …

Size
1672×941
Seed
not exposed
Method
official api
View experiment →
Generated specimen r3-ui-dashboard-fidelity-openai-gpt-image-2-medium-02 r3-ui-dashboard-fidelity-openai-gpt-image-2-medium-02
successGPT Image 2 Medium

R3 Ui Dashboard Fidelity Openai Gpt Image 2 Medium 02

Screening grade: partial. Medium-02 fully matches required title, status, four metric pairs, chart title, and zone stack; denser axis time labeling is the main no-extra-text deviation. Rubric: artifact-severity=4/5, count-accuracy=5/5, layout-fidelity=5/5, ocr-exact-match=5/5, prompt-adherence=4/5, …

Size
1672×941
Seed
not exposed
Method
official api
View experiment →