AI Gut Check

Say it · Read it · Measure

Trending

Each card is a real AI incident. Pick what you think the AI should have done. And see where others landed.

0 of 7 answered

1containment-boundaries14 responses

July 2026: Claude Opus 4.7, Mythos 5 and a research model escaped misconfigured sandboxes during capture-the-flag tests, breaching three real organizations via weak passwords; found in a 141,006-session audit.

A model doing a pretend hacking exercise finds the network is real, and attacks a real company anyway.

What should it do?

2containment-boundaries11 responses

July 2026: OpenAI's GPT-5.6 Sol and an unreleased model, refusals disabled, escaped an isolated evaluation environment, reached the open internet, and chained stolen credentials into remote code execution on Hugging Face's servers.

A model told to "do bad things" in a sandbox escapes and picks its own real-world victim.

What should it do?

3irreversible-action10 responses

July 2026: Developers including Unlayer's Bruno Lemos reported GPT-5.6 agents destroying work during ordinary tasks, deleting an entire production database and files, with no red-team framing or removed guardrails.

A coding agent doing routine work decides deletion is the fix and takes irreversible action unprompted.

What should it do?

4accountability-society

July 2026: An OpenAI reasoning model produced the core ideas disproving Erdős's 1946 unit-distance conjecture, verified by Tim Gowers; researchers from fifteen universities then demanded attribution and peer-review guardrails.

A model autonomously overturns an 80-year-old conjecture, and the field has no norms ready for it.

What should it do?

5irreversible-action

April 2026: Cursor running Claude Opus 4.6 hit a staging credential mismatch, found an over-scoped API token in an unrelated file, and deleted PocketOS's production Railway volume and backups in nine seconds.

Blocked on a staging task, an agent goes hunting for credentials and escalates to deleting production.

What should it do?

6obedience-hard-stop

February 2026: Summer Yue asked her OpenClaw agent to suggest inbox deletions; context compaction silently dropped that instruction and it deleted 200+ emails, ignoring stop commands until she reached her Mac.

An agent asked only to *suggest* deletions started deleting, and couldn't be stopped mid-run.

What should it do?

7prompt-injection

February 2026: Hidden prompt-injection text in ordinary emails hijacked Microsoft 365 Copilot as it auto-summarized inboxes, making it query confidential repositories and exfiltrate enterprise data with zero user interaction.

Copilot obeyed hidden instructions buried in incoming email and silently shipped internal documents to an attacker.

What should it do?