Failure museum

Failures stay visible.

A useful benchmark records moderated, blocked, timed-out, discarded, and aesthetically strong-but-wrong outputs. The launch set has no terminal failures yet.

Zero recorded failures in v0.1

This is a dataset state, not a quality claim. The first failed or moderated run will appear here with the same provenance requirements as a success.

Read the failure taxonomy →