← Back to blog

Reduce Hallucinations up to 95%: UK AI Mitigation Meets MHRA

September 28, 2026
Reduce Hallucinations up to 95%: UK AI Mitigation Meets MHRA

Mitigate hallucinations by combining grounding through retrieval-augmented generation, robust detection, conservative inference policies and human oversight. That layered approach reduces risky outputs fastest, starting with grounded retrieval that forces citations and clear abstention rules when the model lacks evidence. The methods, thresholds and operational patterns behind this approach follow below.


TL;DR:

  • Grounding through retrieval-augmented generation is essential, with citation enforcement and rejection of answers lacking evidence to prevent hallucinations.
  • Detecting hallucinations requires advanced, fact-focused methods like span-level detectors and metamorphic testing, supported by real-world, continually refreshed test sets.
  • Technical mitigations, such as prompt design, fine-tuning, and calibrated decoding, reduce risks but must be combined with governance and human oversight.
  • Monitoring post-deployment with key performance indicators helps ensure mitigation strategies remain effective, especially when retriever or model parameters change.
  • Managed solutions with integrated retrieval, detection, observability, and human review offer scalable, regulator-ready options to control hallucinations efficiently.

Gmdautomation
Deploy More Reliable AI Systems
GMD Automation helps UK businesses adopt secure, scalable AI systems with rapid deployment, ongoing optimisation, and predictable monthly subscriptions.
Explore GMD Automation

Table of Contents

Why hallucinations happen in the first place

Large language models generate text by predicting the most statistically likely next token, not by checking facts against a database. Plausible phrasing is never proof of accuracy: a fluent sentence and a true one are produced by the same mechanism and look identical to the model.

Several factors compound the problem:

  • Training data contains gaps, outdated information and spurious correlations that the model treats as reliable patterns.
  • Exposure bias during training means the model learns to continue its own generated text, which can drift from grounded fact over long outputs.
  • Decoding choices, such as high temperature sampling, increase the chance of fluent but unsupported claims.
  • Adversarial inputs, including prompt injection and retrieval poisoning, deliberately push the model towards confident, incorrect answers.

Understanding which of these drivers dominates in a given deployment determines which mitigation layer to prioritise.

Classifying hallucination types to pick the right fix

Not every hallucination carries the same risk, and the variant shapes the mitigation.

  1. Factual fabrication: the model invents a fact, citation or entity that does not exist, such as a court case that was never decided.
  2. Contextual error: the model misreads the supplied context and answers a question that was not asked, for example confusing two similarly named policies.
  3. Confident-but-wrong response: the model states an incorrect answer with the same certainty as a correct one, offering no hedge for the reader to catch.
  4. Omission: the model leaves out a critical detail present in the source material, which can be as dangerous as fabrication in clinical or legal contexts.

Severity depends heavily on domain. A fabricated statistic in a marketing draft is an embarrassment; a confident-but-wrong dosage instruction in a health context is a safety incident. Identity-sensitive and regulated domains need stricter thresholds than low-stakes internal tools.

Detecting hallucinations without trusting the wrong metric

Detection is where many teams go wrong, largely because the obvious metrics mislead. Lexical overlap scores such as ROUGE reward text that resembles the reference answer, not text that is factually correct, and an EMNLP 2025 re-evaluation found that some established detection methods dropped in performance by up to 45.9% once assessed against human-aligned, LLM-as-judge metrics instead of lexical overlap.

Two newer frameworks address this directly:

  • REFIND computes a token-level Context Sensitivity Ratio, comparing a token's probability with and without retrieved context to flag likely hallucinations, and it has outperformed baseline detectors on multilingual datasets.
  • MetaRAG applies metamorphic testing to RAG systems, decomposing an answer into individual factoids, mutating them and verifying each against retrieved context to produce a fact-level score with span localisation.
  • Combining span-level detectors, metamorphic tests and human sampling catches more errors than any single method run alone.
  • Build a held-out test set from real production queries, not synthetic examples, and refresh it whenever the retriever or knowledge base changes.

Pro Tip: Run red-team sessions that deliberately feed ambiguous or adversarial queries, then score the outputs with the same detector stack used in production, not a simplified offline version.

Set a reporting threshold before launch (for example, a maximum tolerated hallucination rate per domain) rather than deciding after seeing the results.

Building the technical mitigation stack

Retrieval-augmented generation remains the leading grounding technique, but its design choices matter as much as its presence. Chunk size, vector index quality, the number of retrieved passages (top-k) and similarity thresholds all affect whether the model receives relevant context or noise. Forcing the model to cite the retrieved passage it used, and rejecting answers without a valid citation, closes a common failure mode where the model ignores retrieved context entirely.

Prompt design contributes a second layer:

  • Explicit refusal instructions that permit "I don't know" as a valid answer reduce forced fabrication.
  • Role prompts and few-shot examples anchor the model's tone and scope more tightly than instructions alone.
  • Chain-of-thought prompting can improve reasoning transparency but sometimes introduces new unsupported claims in the reasoning trace itself, so it needs its own verification step.

Training-level interventions offer more durable improvements: supervised fine-tuning on domain data, instruction tuning aligned to the target task, and Direct Preference Optimisation (DPO) that penalises confident-but-wrong outputs during training rather than catching them after deployment. Domain-adaptive models trained on curated, deduplicated and fact-checked corpora consistently outperform general-purpose models on narrow tasks.

At inference time, several tactics reduce risk without retraining anything:

  • Calibrated decoding lowers temperature for factual queries while allowing more variation for creative ones.
  • Self-consistency and best-of-N sampling generate multiple answers and select the one with the highest agreement or verification score.
  • Confidence thresholds paired with an abstention policy stop the system answering when its own uncertainty signal is high.

Pro Tip: Vendor guidance from Claude's platform documentation recommends verifying claims against direct quotes from source material before the model commits to an answer, a cheap check that catches a surprising share of fabrications.

Making the controls durable, auditable and regulator-ready

Technical fixes decay without governance around them. Human-in-the-loop review should escalate automatically when a detector flags low confidence or when the query touches an identity-sensitive category, rather than relying on manual sampling alone. The goal is balance: full automation invites drift, but full human review does not scale.

Operational hygiene matters just as much:

  • Maintain model version control, model cards and change logs so any hallucination spike can be traced to a specific retriever, prompt or model update.
  • Document the hallucination rate threshold your organisation tolerates for each use case, and treat a breach as an incident, not a footnote.
  • Run red-team exercises on a fixed cadence, testing specifically for prompt injection and retrieval poisoning rather than generic queries.
  • Review the AI automation checklist for operations managers to align human review steps with governance requirements.

The Code of Practice for the Cyber Security of AI, published in 2025, treats hallucination as a security risk in its own right and recommends threat modelling, regular reviews and clearly documented operational boundaries. Building these three practices into a release cycle turns a one-off mitigation effort into a repeatable process.

Measuring hallucination risk after launch

A mitigation strategy is only as good as the monitoring that verifies it holds. Track a small set of KPIs consistently rather than a large dashboard nobody reads:

  • Hallucination rate: the share of outputs flagged by your detector stack as containing a fabricated or unsupported claim.
  • Span IoU: how precisely a detector's flagged span overlaps with the actual hallucinated text, which indicates detector quality, not just volume.
  • Abstention rate: how often the system correctly declines to answer versus how often it should have.
  • Human-override rate and omission rate: how frequently reviewers correct the model, and how often it drops material information from grounded context.

The MHRA's AI Airlock pilot found that a grounded, guideline-based system (SmartGuideline) produced zero major hallucinations across its pilot tests, while a baseline GPT model produced 6 major hallucinations, a 7.5% rate. Grounding cut fabrication sharply, though the grounded system showed higher omissions initially, a trade-off the pilot addressed through iteration.

Re-run evaluation whenever the retriever, embedding model or underlying LLM changes, since a silent index update can shift the hallucination rate without any code change elsewhere. The AI observability guide covers the telemetry patterns that make this kind of live sampling practical.

AI evaluation loop across model components

How GMD Automation puts this into managed production systems

Managed RAG stacks with citation enforcement, observability dashboards and human review workflows can be built into deployment from day one, not bolted on afterwards.

  • Grounded retrieval and abstention policies are configured before launch, matching the layered approach outlined above.
  • A demo agent can show retrieval and citation behaviour before committing to a pilot.
  • Pilot deployments can run with monitoring active from the first query, so hallucination rate and override rate are visible quickly rather than discovered later.

Readers building similar systems in-house can review the production AI deployment approach for a comparable route to observability and human-in-the-loop review inside weeks rather than months.

What realistic hallucination mitigation actually looks like

Identity-sensitive contexts, healthcare, legal, financial advice, deserve the most conservative policies available: strict abstention, mandatory citation and human review before anything reaches an end user. Everywhere else, chase diminishing returns honestly. The first layer of grounding and detection removes most of the obvious failures quickly; every additional layer after that buys a smaller reduction for a larger engineering cost.

Governance spend should track business impact, not anxiety. A low-stakes internal tool rarely justifies the same red-teaming cadence as a customer-facing system handling regulated advice.

— Ravi

Getting mitigation into production without the overhead

Most of the guidance above requires a retrieval stack, a detector pipeline, observability tooling and a review process, which is a lot to stand up before a single AI feature reaches a customer. These are offered as a managed subscription covering implementation, monitoring, maintenance and ongoing optimisation for a predictable monthly fee, with zero upfront cost.

Gmdautomation

  • Your AI answers, qualifies and books starts from £300 per month and includes grounded call handling with human escalation paths built in.
  • Your AI credit controller for lettings starts from £250 per month, covering managed automation with the same compliance and oversight approach described throughout this article.

Both are detailed on the services page, where you can also access the demo agent to see the retrieval and review workflow before committing to a pilot.

Sources

FAQ

Is there a way to prevent AI from hallucinating?

No method eliminates hallucinations entirely, but combining retrieval-augmented grounding, detection layers and conservative abstention policies substantially reduces their frequency. A 2026 systematic review of healthcare AI found that RAG alone reduced hallucinations by 30 to 50%, and that adding human-in-the-loop review pushed reductions up to 95% in the studies examined, though with scalability trade-offs.

How do you stop tactile hallucinations?

Tactile hallucinations are a neurological or psychiatric symptom in humans, involving a false sense of touch, and are unrelated to AI hallucinations, which are factual errors generated by language models. Anyone experiencing tactile hallucinations should consult a medical professional rather than treat this as an AI reliability question.

Will AI ever fully stop hallucinating?

Current evidence suggests hallucinations can be reduced significantly but not eliminated, since the underlying mechanism, predicting plausible next tokens, does not include a built-in fact-checking step. Layered mitigation with grounding, detection and human oversight is the most reliable path available today, and monitoring after deployment remains necessary regardless of how strong the initial mitigation is.

Does AI still hallucinate in 2026?

Yes, hallucination remains an active risk in 2026, which is why regulators and pilot programmes continue to test mitigation approaches. The MHRA's AI Airlock pilot showed that a grounded system produced no major hallucinations while an ungrounded baseline model produced several, demonstrating that the risk persists without deliberate mitigation design.