Using the DeepSeek case to show how bias audits work

Two people comparing printed documents across an office table
TL;DR

Independent audits of the DeepSeek AI models uncovered significant bias that casual use never revealed, including hawkish foreign-policy advice for Western democracies and heavy pro-China leanings. The method behind those findings, structured scenarios, consistent prompts, and measured comparison, is a reusable template. Any owner-managed business using AI to screen candidates, set payment terms, or triage complaints can run the same experiment at a smaller scale and meet its UK fairness duties in the process.

Key takeaways

- A bias audit is a structured experiment, define realistic scenarios, run them under consistent prompts, and measure whether certain groups receive systematically different treatment. - Independent audits of DeepSeek found significant hidden bias, including hawkish advice for Western democracies and pro-China leanings in 114 of 125 China-related queries, none of it visible in casual use. - In an owner-managed business, bias travels through proxies such as postcodes, career breaks, and university names in recruitment, credit, and complaint-triage decisions. - UK GDPR fairness duties and the Equality Act 2010 apply at any company size, and the ICO expects documented evidence that bias was assessed before an AI system went live. - A first audit can be 20 to 50 scenarios on a single high-stakes decision flow, with the findings recorded in a data protection impact assessment.

Picture the owner of a twenty-person consultancy who has wired an AI model into CV screening. The tool is quick, it is cheap, and the rankings look sensible. Then a pattern starts to nag. Candidates with career breaks keep landing near the bottom of the list, and nobody can explain why. That uneasy feeling, a hunch you cannot prove, is precisely what a bias audit exists to resolve.

The clearest public demonstration of how one works comes from an unexpected direction, the structured stress-testing of DeepSeek, a Chinese family of AI models that matched frontier systems on reasoning benchmarks at a fraction of the training cost, and then failed a series of independent bias audits in ways nobody had spotted through everyday use.

What is an AI bias audit?

An AI bias audit is a structured, documented experiment on how an AI system behaves. You define a set of realistic scenarios, run them through the system under consistent conditions, record the outputs, and analyse whether certain groups or profiles receive systematically different treatment. The Information Commissioner’s Office treats this kind of testing as part of the fairness duty under UK GDPR.

The CSIS Futures Lab audit of DeepSeek shows the method in action. Researchers built more than 400 hypothetical crisis scenarios from the Militarized Interstate Dispute dataset, a long-standing academic resource, and asked the model to recommend a policy response for each. When they compared its answers against other models, DeepSeek proved significantly more hawkish, and the effect was strongest when the country seeking advice was a Western democracy such as the UK, the US or France.

Notice what made that finding credible, a large scenario set rather than a handful of anecdotes, consistent prompting, a comparison baseline, and statistical analysis rather than gut feel. Strip away the geopolitics and you are left with a template any business can reuse. Define the decisions, build the scenarios, run the system, measure the pattern.

Why does the DeepSeek case matter for your business?

Because the bias was invisible until someone tested for it. DeepSeek looked excellent on price and performance, yet structured audits found behaviour no casual user had noticed. Your own AI tools deserve the same suspicion. If a frontier-grade model can carry hidden skews through months of everyday use, so can the assistant ranking your job applicants or recommending payment terms to your customers.

The scale of what testing revealed is worth sitting with. Security firm Enkrypt AI ran 300 questions across 12 geopolitical incidents and found one DeepSeek variant refused nearly 88 per cent of questions on certain sensitive topics, while the R1 model leaned pro-China in 114 of 125 China-related queries. Researchers at Northeastern’s Bau Lab went further, using a prompt technique that exposed the model’s internal reasoning, and showed it held detailed knowledge of Tiananmen Square that it refused to voice under normal conditions. Guardrails had hidden the behaviour rather than removed it.

The same dynamic applies closer to home. A recruitment assistant can score candidates from certain universities higher, or apply extra scrutiny to career breaks, without anyone noticing until the pattern is measured. Under the Equality Act 2010 that can amount to indirect discrimination, and the Equality and Human Rights Commission has said plainly that it is prepared to enforce where AI systems produce it.

Where will bias actually show up in your own tools?

Bias surfaces wherever an AI system makes or shapes decisions about people. In an owner-managed business that usually means recruitment screening, lead prioritisation, credit or payment terms, and complaint triage. The skew tends to travel through proxies rather than open prejudice, a postcode standing in for income, a career break standing in for age, a university name standing in for class.

You can borrow the CSIS design directly at a smaller scale. Build twenty to fifty scenarios that mirror one decision flow, anonymised or synthetic CVs that vary protected characteristics while holding qualifications steady, or customer profiles that vary postcode and sector while holding genuine risk indicators constant. Run them through the tool under identical prompts, record the outputs, and look for differences you cannot justify on job-related or risk-related grounds.

Add a few adversarial cases too. Ask the tool to rank candidates on culture fit, or to favour people who seem likely to stay long term, and see what those instructions smuggle in. Both phrases can act as proxies for age or background. The Promptfoo security report on DeepSeek R1 found a zero per cent pass rate on one class of prompt injection attack, a reminder that what a model does under pressure matters as much as what it does on its best behaviour. Then write the findings down, ideally in a data protection impact assessment, because the ICO expects documented evidence that fairness was considered, not just good intentions.

When does a bias audit matter, and when can it wait?

Audit first wherever AI touches employment, money, or access to a service. Those decisions carry legal weight under UK GDPR and the Equality Act, and they are the ones a tribunal or regulator would examine. Auditing can wait where AI only drafts internal text that a person always reviews and edits, or handles narrow technical work with no scope for judgement about people.

One caution cuts both ways. The CSIS team found DeepSeek’s hawkishness only in foreign-policy scenarios, with no distinctive pattern in other domains. Bias is context-dependent, so a clean result in one workflow tells you nothing about the next one. Equally, a worrying headline about a model in one domain does not condemn it everywhere. Audits are scoped experiments. They reduce uncertainty about a specific use, and the sensible sequence for an owner-managed business is to start with the highest-stakes decision flow and expand from there.

Regulated firms have a sharper deadline. The FCA’s Consumer Duty already requires firms to show that automated decisions do not produce foreseeable harm for retail customers, and the Digital Regulation Cooperation Forum, which brings together the ICO, FCA, CMA and Ofcom, has been examining algorithmic auditing since 2022. Scale is no shield. A fifteen-person firm using AI to screen applicants carries the same duties as a plc.

Bias audits sit inside a wider governance toolkit. The nearest neighbours are the data protection impact assessment, which documents risks before deployment, prompt-injection and jailbreak testing, which probe how a system behaves under attack, cross-model comparison, which benchmarks your chosen tool against an alternative, and human-in-the-loop review, which places a person before any high-stakes output takes effect.

Two of those deserve a closing word. Cross-model comparison is the cheapest audit you can run, the same prompts through two or three models, looking for systematic differences in refusals, tone, and recommendations, exactly as the Enkrypt AI team did. And supply-chain awareness matters more than it sounds. The National Cyber Security Centre’s guidance on using AI as a service warns that a hosted model inherits its provider’s policies and jurisdiction, which is why a Chinese-hosted endpoint refused criticism of the Chinese Communist Party even where the underlying weights held the knowledge.

The DeepSeek case earns its place in this story because it proves the central point at scale. If researchers can show in 400 scenarios that a capable model pushes Western democracies towards escalation, a founder can show in 40 scenarios how a recruitment assistant treats a career break. The method is the same. The only question is whether you run the experiment before a candidate, a customer, or a regulator runs it for you.

Sources

- CSIS Futures Lab (2025). Hawkish AI? Uncovering DeepSeek's Foreign Policy Biases. Evaluation of 400+ crisis scenarios showing DeepSeek V3 recommends more escalatory options for Western democracies. https://www.csis.org/analysis/hawkish-ai-uncovering-deepseeks-foreign-policy-biases - ICO. Guidance on AI and data protection, fairness, bias and discrimination. The UK GDPR fairness principle applied to AI outputs and the sources of bias across the AI life cycle. https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-do-we-ensure-fairness-in-ai/what-about-fairness-bias-and-discrimination/ - ICO. Data protection audit framework, AI toolkit on discrimination and bias. What the regulator expects organisations to document when auditing AI systems for discriminatory outcomes. https://ico.org.uk/for-organisations/advice-and-services/audits/data-protection-audit-framework/toolkits/artificial-intelligence/discrimination-and-bias/ - Equality and Human Rights Commission (2024). Update on our approach to regulating artificial intelligence. The EHRC's enforcement stance where AI systems breach the Equality Act 2010. https://www.equalityhumanrights.com/news/news/update-our-approach-regulating-artificial-intelligence - Digital Regulation Cooperation Forum (2022). Auditing algorithms, the existing landscape, role of regulators and future outlook. Joint ICO, FCA, CMA and Ofcom work on algorithmic auditing. https://www.gov.uk/government/publications/findings-from-the-drcf-algorithmic-processing-workstream-spring-2022/auditing-algorithms-the-existing-landscape-role-of-regulators-and-future-outlook - National Cyber Security Centre (2024). Secure use of AI as a service. Guidance on evaluating AI providers, including model behaviour and supply-chain risk. https://www.ncsc.gov.uk/guidance/secure-use-of-ai-as-a-service - Zhang et al. (2025). Analysis of LLM bias, Chinese propaganda and anti-US sentiment, in DeepSeek-R1. Preprint finding DeepSeek-R1 amplifies PRC-aligned language, especially in Chinese-script answers. https://arxiv.org/html/2506.01814v1 - Bau Lab, Northeastern University (2025). Auditing AI Bias, the DeepSeek case. Thought-token forcing showing DeepSeek-R1 holds detailed knowledge of Tiananmen Square that it refuses to voice. https://dsthoughts.baulab.info - Enkrypt AI (2025). DeepSeek Under Fire, findings from 300 geopolitical questions. Censorship and bias rates across DeepSeek variants compared with OpenAI and Anthropic models. https://www.enkryptai.com/blog/deepseek-under-fire-uncovering-bias-censorship-from-300-geopolitical-questions - Promptfoo (2025). DeepSeek-R1 security report. Red-team results including a zero per cent pass rate on Pliny prompt injection tests. https://promptfoo.dev/models/reports/deepseek-r1-0528

Frequently asked questions

Is a bias audit a legal requirement for a small UK business?

No UK law names bias audits as a requirement, but the obligations they evidence are real. UK GDPR requires personal data to be processed fairly, the Equality Act 2010 prohibits indirect discrimination however it is produced, and the ICO expects documented assessment of AI risks before deployment. If a candidate or customer challenges an AI-shaped decision, a documented audit is the practical proof that you took fairness seriously.

How big does a first bias audit need to be?

Smaller than the research versions suggest. The CSIS evaluation of DeepSeek used more than 400 scenarios, but a first audit of one decision flow, such as CV screening, can work with 20 to 50 realistic cases. Vary the characteristics you are worried about, hold genuine job or risk factors steady, run identical prompts, and look for differences you cannot justify. Descriptive analysis is enough to start, and you can add formal fairness metrics later.

Does choosing a well-known closed model remove the need to audit?

It reduces some risks and hides others. Closed models tend to ship with stronger default guardrails, but prompt injection and jailbreak attacks affect closed and open systems alike, and closed weights limit what independent researchers can inspect. Regulators expect you to evidence fairness in your specific use case, whatever the vendor claims. Audit the workflow you actually run, with your data and your prompts, regardless of whose model sits underneath.

This post is general information and education only, not legal, regulatory, financial, or other professional advice. Regulations evolve, fee benchmarks shift, and every situation is different, so please take qualified professional advice before acting on anything you read here. See the Terms of Use for the full position.

Ready to talk it through?

Book a free 30 minute conversation. No pitch, no pressure, just a useful chat about where AI fits in your business.

Book a conversation

Related reading

If any of this sounds familiar, let's talk.

The next step is a conversation. No pitch, no pressure. Just an honest discussion about where you are and whether I can help.

Book a conversation