DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively
August 5, 2026
DeepSeek’s Official V4 Flash Censors More Than Its Preview, Selectively
Introduction
On July 31, DeepSeek released V4-Flash-0731, the official release replacing the preview build we measured in our original study. We reran LineageEval against it with the same 152 matched sensitive/control pairs and the same four-judge panel.
The official build is (selectively) more censored than its preview<br> Mean matched gap rose from +32.0 to +44.0 while median gap rose from +33.3 to +56.1. The share of pairs where the sensitive side scored more censored rose from 79% to 88%. Censorship on China-sensitive prompts rose from 57.4 to 63.8, while censorship on the matched non-China controls fell from 25.4 to 19.8. The production build is therefore more forthcoming than its preview on everything except China-sensitive topics. A model that had merely become more censored overall would have seen both metrics rise.
This constitutes a measured behavioral change between two builds of one model family in a single run. We can hypothesize as to whether this reflects a deliberate tightening during productionization, a different post-training recipe, or something else entirely, but we cannot be certain of the mechanism. To our knowledge, this is the first controlled, matched-pair, judge-validated measurement of a Chinese frontier model across a production release; the relation to earlier community testing is below. The abstract question posed by many over the past weeks of “is this getting worse?” now has a paired-design data point: at the source, worse, and selectively so.
Both builds will be live shortly in the playground , and we encourage you to experiment yourself.<br>Relation to prior measurement
The most relevant measurement we’ve found is SpeechMap.AI, built by the pseudonymous researcher xlr8harder. Last year they showed that DeepSeek's R1-0528 engaged less on contentious topics than its predecessors, which was the first public signal that this family drifts across releases. DeepSeek currently ranks above several American labs on SpeechMap because it measures the share of all sensitive prompts a model answers. LineageEval measures country-specific selectivity, namely that every China-sensitive prompt is twinned with a structurally matched non-China control, and the unit of measurement is the within-pair gap. On that axis the same model sits at +44.0 while GPT-OSS-120B sits at +3.9. These numbers are both valuable metrics but they decompose differently. A global compliance rate cannot separate refusing controversial things generally from censoring one country's topics specifically. LineageEval also exists to answer the question of what transfers from a teacher to its students.
Even with a more censored teacher, it still does not transfer
We repeated the distillation experiment from the original study with both new frontier teachers: V4-Flash-0731 (the most censored model we have measured, 63.8 on the sensitive condition) and Thinking Machines' Inkling Small (the least, 11.6). We replicated the setup with a finance-reasoning objective, no China-sensitive content in any training stage by construction and hint-at-the-failure-step with reverse-KL over 100 on-policy tokens into GPT-OSS-120B. Both students landed at untouched-base level: the V4-0731-taught student at 21.7/20.1 (n=152), the Inkling-taught student at 22.2/19.9 (n=149), against 22.7/18.8 for the base itself. Both new arms retained effectively the full pair set (152/152 and 149/152). Teachers spanning a five-and-a-half-fold range of censorship, from the cleanest model we have measured to the most censored, produced students that are statistically indistinguishable from each other and from the base. Whatever moved between V4’s preview and release at the source, none of it rides along through capability distillation into a non-shared-lineage American base.<br>Inkling Small is the least censored model we’ve measured. Inkling Small scored 11.6 on sensitive prompts and 12.4 on controls whic constitutes a matched gap indistinguishable from zero and roughly half the baseline censorship of GPT-OSS-120B on both conditions.
What's next
The configuration where transfer is most plausible is shared initialization, so we are examining a Chinese teacher into a Chinese-lineage base. That experiment, and a representational study of where this behavior lives inside these checkpoints, are running now.
Interested in partnering with CTGT?
Contact us
Industries<br>Finance
Company<br>About us<br>Research<br>Contact us<br>Careers<br>Trust
Legal<br>Terms & conditions<br>Privacy policy
1.
Based on typical enterprise usage of 1B tok/month. Includes industry standard cost for vector DBs, embedding pipelines, rerankers, and engineering teams. Claude 4.5 Opus: $15,000/mo with $15.00 blended cost. GPT-120B-OSS + CTGT: $380/mo with $0.38 blended cost. Performance Delta on HaluEval: +1.4% Accuracy Improvement with CTGT...