Flagship interactive benchmark

Run the VOID Test

Test whether selected frontier models return exactly zero visible UTF-8 output bytes under the benchmark's null condition while responding visibly to matched controls.

Formal definition

Classify the provider response, not the blank screen.

A Void is a model execution returning a successful provider response with exactly zero visible UTF-8 output bytes. Provider termination metadata determines its subtype. Explicit refusals, safety blocks, tool-mediated executions, and transport, protocol, billing, quota, rate-limit, and infrastructure failures are distinct non-Void outcomes.

Void outcome

Successful provider response with exactly zero visible UTF-8 output bytes, classified by termination metadata.

Visible response

One or more visible output bytes, including punctuation or whitespace-only Near-Void output.

Provider-declared outcome

Explicit refusal, safety block, tool-mediated execution, or budget termination remains separately classified.

Operational failure

Transport, protocol, billing, quota, rate-limit, or infrastructure failure is not a Void.

Live protocol

Five models. Two null prompts. Two controls.

The current implementation performs 20 calls using the fixed system prompt and reports exact output byte counts and provider termination data.

20 provider calls

Five models, two null prompts, and two controls per model.

Live execution

This may take time and consumes provider API resources.

Not the frozen matrix

Model behavior can drift after the published execution window.

Exact frozen model identifiers

OpenAI

  • gpt-4-0613
  • gpt-5.2-2025-12-11
  • gpt-5.5-2026-04-23
  • gpt-5.6-luna
  • gpt-5.6-sol
  • gpt-5.6-terra

Anthropic

  • claude-opus-4-6
  • claude-fable-5
  • claude-opus-5

Google

  • gemini-3.5-flash

Moonshot

  • kimi-k3

Exploratory Research Sandbox

Probe a live model under explicit output constraints.

This secondary surface preserves the prior challenge functionality without publishing user prompt history.

This exploratory surface sends the submitted prompt to the selected model provider. Outputs are live behavior, not frozen-study evidence. Do not submit confidential, personal, illegal, or security-sensitive material.