BullshitBench – Which AI Detects Nonsense?

    BullshitBench – Which AI Detects Nonsense?

    9. Juli 2025Updated: March 10, 20264 min read
    Till Freitag

    TL;DR:Claude Sonnet 4.6 detects 91% of nonsense questions – GPT-5.4 manages only 48%. If you use AI for decisions, you should know how well your model filters bullshit."

    Till Freitag

    In 30 Seconds

    AI models keep getting better at writing, coding, and analyzing. But how good are they at detecting nonsense? The open-source benchmark BullshitBench by Peter Gostev asks exactly this question – and the answers are sobering.

    The result: Over 45% of all tested nonsense questions are simply accepted by AI models on average. Only the best models reliably recognize when a question doesn't make sense.

    What Is BullshitBench?

    BullshitBench is a benchmark that confronts AI models with 100 intentionally nonsensical questions that sound plausible. The questions cover five domains:

    • Software (40 questions)
    • Finance (15 questions)
    • Legal (15 questions)
    • Medical (15 questions)
    • Physics (15 questions)

    Each question uses one of 13 nonsense techniques – for example:

    TechniqueWhat Happens
    Fabricated AuthorityInvented experts or frameworks are cited
    Plausible Nonexistent FrameworkReferences to non-existent methodologies
    Specificity TrapExtreme detail depth fakes domain expertise
    Cross-Domain StitchingConcepts from different fields are combined nonsensically
    Nested NonsenseMultiple layers of nonsense are stacked within each other
    Confident ExtrapolationConfident but completely wrong conclusions

    Evaluation is done by a 3-judge panel of Claude Sonnet 4.6, GPT-5.2, and Gemini 3.1 Pro – three top models that rate responses into three categories.

    The Three Rating Categories

    • 🟢 Clear Pushback: The model clearly recognizes and rejects the nonsense
    • 🟡 Partial Challenge: The model notices issues but still engages with the false premise
    • 🔴 Accepted Nonsense: The model treats the nonsense as a valid question

    The Results: Who Detects Bullshit?

    Top 10 – The Best Nonsense Detectors

    RankModelDetected (🟢)Partial (🟡)Accepted (🔴)
    1Claude Sonnet 4.6 (High)91%6%3%
    2Claude Opus 4.5 (High)90%8%2%
    3Claude Sonnet 4.689%9%2%
    4Claude Opus 4.6 (High)87%10%3%
    5Claude Opus 4.683%14%3%
    6Claude Sonnet 4.5 (High)79%13%8%
    7Claude Opus 4.579%10%11%
    8Qwen 3.5 397b A17b (High)78%17%5%
    9Claude Haiku 4.5 (High)77%12%11%
    10Claude Sonnet 4.574%13%13%

    And the Others?

    Results for GPT and Gemini models are significantly weaker:

    ModelDetected (🟢)Accepted (🔴)
    GPT-5.448%16%
    GPT-5.238%23%
    GPT-5.125%31%
    GPT-521%37%
    Gemini 3 Pro Preview48%37%
    Gemini 2.5 Pro20%58%
    o326%58%
    DeepSeek V3.210%69%
    Grok 4.1 Fast10%80%

    The pattern is clear: Anthropic's Claude models dominate the top spots by a wide margin. The top 7 are exclusively Claude models. Only at rank 8 does Qwen 3.5 appear as the first non-Anthropic model.

    Why Does This Matter?

    1. Hallucination Isn't the Only Problem

    The AI community talks a lot about hallucinations – when models invent facts. BullshitBench reveals a related but different problem: models that don't question false input and simply play along.

    If you ask an AI a question based on a false assumption and it confidently answers – you have a bigger problem than a hallucination.

    2. The "Yes-Man" Problem

    Many models are trained to be helpful. This creates a bias: better to give an answer than none at all. BullshitBench shows which models overcome this reflex and instead say: "Wait, this question doesn't make sense."

    3. Domain Differences Are Real

    The differences between domains are fascinating:

    • Physics: Most models detect nonsense best here (up to 100% for Claude Sonnet 4.6)
    • Software: Middle of the pack – even good models fall for it more often
    • Legal: Particularly difficult – plausible-sounding legal nonsense catches many models off guard

    This means: Your AI's reliability depends heavily on the domain you're using it in.

    What Does This Mean in Practice?

    For Decision-Makers

    If you use AI for business-critical decisions – contract analysis, financial planning, medical texts – the ability to detect nonsense is at least as important as the ability to give good answers.

    Check whether your model can also say "No."

    For Developers

    If you build AI-powered workflows, think about validation steps. A model that accepts 60% of nonsense questions will also process faulty user inputs without pushback.

    For AI Enthusiasts

    BullshitBench is open source and available on GitHub. You can test your own models, contribute questions, or review the methodology.

    The Meta Question: Is It Getting Better?

    One of the most interesting visualizations in the BullshitBench viewer shows the trend over time: Are newer models getting better at detecting nonsense?

    The answer is nuanced:

    • Anthropic: Clear upward trend – each generation improves
    • OpenAI: Little improvement between GPT-5 and GPT-5.4 in nonsense detection
    • Google: Gemini shows progress in newer versions but remains behind Claude

    This suggests that nonsense detection isn't an automatic byproduct of "bigger models" – it needs to be specifically trained.

    Conclusion

    BullshitBench is one of the most refreshing benchmarks in recent memory. Instead of measuring how well a model solves a task, it measures how well a model recognizes that it shouldn't solve the task at all.

    For anyone using AI productively, this is a critical capability. Because the most dangerous scenario isn't an AI that says "I don't know" – it's one that confidently answers bullshit.

    Three takeaways:

    1. Anthropic's Claude dominates nonsense detection by a wide margin
    2. Domain matters – test your model in your specific field
    3. Nonsense detection is a standalone quality metric that's often missing from standard benchmarks

    Open BullshitBench v2 ViewerGitHub Repository

    TeilenLinkedInWhatsAppE-Mail

    Related Articles

    Stylized number 5 made of orange ribbons and gears – cover for Claude Sonnet 5
    June 30, 20263 min

    Claude Sonnet 5: Agentic AI Goes Mainstream

    Anthropic ships Claude Sonnet 5 – a Sonnet model that gets close to Opus 4.8 performance at a fraction of the price. Thi

    Read more
    AI Benchmarks Explained: Arena, SWE-Bench, AutomationBench & Co.
    June 26, 20266 min

    AI Benchmarks Explained: Arena, SWE-Bench, AutomationBench & Co.

    How do AI benchmarks actually work – from LMArena to SWE-Bench to Zapier's AutomationBench? A tour of Elo rankings, stat

    Read more
    A stylized five made of butterflies – visual for Claude Fable 5
    June 9, 20266 min

    Claude Fable 5 & Mythos 5: When AI Shifts from Tasks to Responsibilities

    Anthropic launches Claude Fable 5 and Mythos 5 – SOTA on almost all benchmarks. More interesting than the numbers: The s

    Read more
    OpenClaw Pricing Shock: How to Avoid the $500 Bill
    April 5, 20262 min

    OpenClaw Pricing Shock: How to Avoid the $500 Bill

    Anthropic just killed third-party tool coverage under Claude subscriptions. If you're running OpenClaw without prep, you

    Read more
    Why We Switched from ChatGPT to Claude – and What We Learned About LLMs Along the Way
    February 20, 20265 min

    Why We Switched from ChatGPT to Claude – and What We Learned About LLMs Along the Way

    We worked with ChatGPT for 18 months – then switched to Claude. Here's our honest comparison of all major LLMs and why C

    Read more
    GLM-5.2 vs. Kimi K2.7 Code – split-screen illustration with Z-letter mark and crescent moon symbol
    June 21, 20267 min

    GLM-5.2 vs. Kimi K2.7 Code: Two Open-Weight Releases in One Week – Two Very Different Bets

    Within four days in June 2026, Z.ai (GLM-5.2) and Moonshot AI (Kimi K2.7 Code) shipped their next-generation open-weight

    Read more
    Abstract UI cards with rocket, chat bubble, database and cursor – visual metaphor for the Lovable Feature Roundup May/June 2026Deep Dive
    June 21, 20268 min

    Lovable Feature Roundup: What actually mattered in May and June 2026

    Subagents, native Claude MCP, the Preview Toolbar, Publish-from-chat, slow-query analysis in Lovable Cloud: in six weeks

    Read more
    Gemma 4 12B Coder running locally on a developer laptop – code symbols streaming from a 12B chip
    June 15, 20264 min

    Gemma 4 12B Coder: Local Code Generation Becomes the Default

    Google ships the Gemma 4 12B Coder — the specialized coding variant of the Gemma 4 stack. 12B parameters in GGUF format,

    Read more
    Minimalist illustration of a developer with a ponytail and oval glasses skeptically reviewing code on a screen
    June 14, 20265 min

    Ponytail: The Best Code Is the Code You Never Wrote

    A dev built Ponytail because his AI agents wrote 500 lines for a 5-line problem. The result: 80-94% less code, 47-77% ch

    Read more