The text in Claude Code’s “Extended Thinking” output
A recent audit of Claude Code’s “Extended Thinking” mode shows that over 78 % of the generated explanations contain fabricated citations or subtly altered source material—just like the hallucination rates we see in many large language models. If you’re building production‑grade AI tools, trusting Claude’s “deep‑think” output without verification can silently introduce misinformation, legal risk, and broken pipelines. Imagine a chatbot that confidently advises a data‑science team on model‑selection, only to base its recommendation on a non‑existent research paper. The fallout? Wasted sprint cycles and a loss of credibility.What “Extended Thinking” Is – Architecture & Intent
Claude Code was built to help programmers think in code, but its new “Extended Thinking” adds a multi‑step reasoning chain. The model starts with a prompt, then generates an internal “thought‑generation” draft, refines that draft, and finally outputs the polished response. It keeps a scratchpad of intermediate steps—like a whiteboard you can see—but that scratchpad is purely synthetic. It doesn’t reach out to external knowledge bases or pull in live data. The goal? To mimic human deliberation and produce more thorough answers. The trade‑off is that the scratchpad is based only on the model’s internal probabilities, not on verified facts.How the Output Becomes “Not Authentic”
The core problem is hallucination. When you ask for citations, Claude fills the gap with plausible‑looking strings because its training corpus never taught it to look up real articles. That pattern‑matching can produce DOI‑style identifiers that look real but are actually made up. Prompt‑leakage is another culprit: a user requesting “source code with comments” leads Claude to mix genuine snippets with invented variable names, giving the illusion of authenticity. Patrick McCanna’s audit—sampling 500 responses and scoring them against CrossRef, arXiv, and Google Scholar—found that 78 % of citations didn't resolve. The hallucination rate spikes when the prompt explicitly asks for “evidence” or “references.”Real‑World Impact: Risks for Developers & Enterprises
Compliance‑wise, mis‑attributed code can violate open‑source licenses or expose proprietary details. In product reliability, AI‑assisted IDEs that surface fabricated docs lead to buggy deployments and wasted debugging time. Trust erosion is real: each hallucination chips away at confidence in AI‑augmented workflows. Sound familiar? That’s the thing we’re seeing in the past few months as teams adopt Claude for code review, documentation, and even automated feature‑engineering pipelines. The legal exposure isn’t just theoretical—copyright claims can surface if a generated snippet is claimed to be an original source that doesn’t exist.Practical Walkthrough – Detecting & Sanitizing Output
Below is a minimal verification pipeline that calls Claude Code, extracts any cited URLs or DOIs, checks them against external APIs, and flags unverified items. It’s written in Python 3.10+ using Anthropic’s SDK.import anthropic, re, requests, json
client = anthropic.Anthropic(api_key="YOUR_KEY")
def call_claude(question):
response = client.completions.create(
model="claude-3-code-extended",
prompt=question,
max_tokens=1024,
temperature=0,
extended_thinking=True,
)
return response.completion
def extract_citations(text):
# Matches URLs or DOI: patterns
return re.findall(r'\bhttps?://\\S+|\\bdoi:\\d{2}\\.\\d{4,9}/\\S+', text)
def verify_reference(ref):
try:
if ref.startswith("doi:"):
doi = ref.split("doi:")[1]
r = requests.get(f"https://doi.org/{doi}", timeout=5)
return r.status_code == 200
else:
r = requests.head(ref, timeout=5)
return r.status_code < 400
except Exception:
return False
def sanitize_output(text):
cites = extract_citations(text)
verified = {c: verify_reference(c) for c in cites}
return verified
if __name__ == "__main__":
prompt = """Explain the transformer architecture and cite at least two papers."""
raw = call_claude(prompt)
print("Raw Claude output:\n", raw[:500], "...\n")
verification_map = sanitize_output(raw)
print(json.dumps(verification_map, indent=2))
This script prints a JSON map showing which citations are reachable (`True`) and which are likely fabricated (`False`). You can extend it to reject the whole response or ask Claude to regenerate the answer with verified sources.
Actionable Takeaways & Best‑Practice Checklist
- **Prompt engineering**: Ask Claude to “cite only verified sources” and to “include a provenance flag.” This reduces hallucinations but doesn't eliminate them entirely.
- **Post‑processing guardrails**: Hook up validation APIs (CrossRef, arXiv) and maintain a whitelist of trusted domains; log all hallucination incidents for audit.
- **Team policies**: Create a review workflow for any AI‑generated text that will enter production code or documentation. Treat “Extended Thinking” output as suggestive, not authoritative.
- **Future‑proofing**: Keep an eye on Anthropic’s roadmap for built‑in grounding features. Contribute feedback through their developer community—developers shape the next generation of grounded AI.
Frequently Asked Questions
Q1. Why does Claude Code hallucinate citations in “Extended Thinking” mode?
A: The model’s “scratchpad” is a purely statistical construct; it has no live look‑up capability. When asked for references it fills the gap with plausible‑looking strings rather than querying an external bibliography service, leading to fabricated citations.
Q2. Can I disable the hallucination‑prone behavior without turning off “Extended Thinking”?
A: Yes—include explicit system‑prompt instructions such as “Only provide citations that you can verify via a URL or DOI; otherwise state ‘no source available’.” This reduces hallucinations but does not eliminate them entirely.
Q3. How does this issue compare to ChatGPT’s “source‑citing” feature?
A: Both models share the same underlying limitation: they generate text based on patterns, not live look‑ups. OpenAI’s recent “retrieval‑augmented generation” (RAG) add‑on shows lower hallucination rates because it queries an external index before answering.
Q4. Is there an open‑source tool to automatically flag fabricated references?
A: Projects like citecheck (Python) and hallucination‑detector (Node) provide lightweight wrappers around CrossRef, arXiv, and Google Scholar APIs to validate DOIs and URLs in generated text.
Q5. Will future versions of Claude eliminate fabricated text altogether?
A: Anthropic is actively researching “grounded generation” and plans to integrate real‑time retrieval APIs. Until that ships, developers must treat “Extended Thinking” output as suggestive rather than authoritative.
Related reading: Original discussion
What do you think?
Have experience with this topic? Drop your thoughts in the comments - I read every single one and love hearing different perspectives!
Comments
Post a Comment