The latest AI safety index just dropped a verdict that sounds damning: Anthropic gets a C+, OpenAI gets a C. Headlines will scream about failing grades. But here is the problem—this scorecard measures public commitments, governance paperwork, and transparency optics. It does not measure model behavior, jailbreak resistance, or actual inference-time harm. Code is the only law that compiles without mercy. And this index doesn't read code.
Let me be clear about what I did not find. I pulled the source material expecting technical specifications—RLHF curves, red-team frameworks, or at least a list of audited checkpoints. None of that exists. What the index actually tracks is a company's willingness to sign pledges, publish 'Responsible Scaling Policies,' and disclose board-level oversight. It is a governance audit. It is not a technical verdict. This is a distinction most coverage fails to make, and it matters because markets, procurement teams, and regulators are starting to treat this C+ like a benchmark for safety.
Context: The Methodology Vacuum
The index under review is a classic 'black-box governance' product. It compiles a score based on public commitments, organizational structure, and the cadence of safety-related blog posts. The score does not weigh a single adversarial attack, a single hallucination log, or a single test of model manipulation. To be blunt: a company could have a flawless technical safety record but a sloppy PR department and still score a C. The reverse is also true—a company with a beautiful compliance document and a catastrophic exploit history could score a B.
We are comparing paper promises, not execution. I have audited smart contracts where the whitepaper promised decentralization but the code showed a kill switch. This feels like the same pattern. The C+ is a measure of public posture, not of actual system safety.
Core: The Governance-Versus-Capability Gap
The real insight here is the stark divorce between 'safety score' and 'technical capability.' Anthropic is a top-tier lab with deep research in interpretability. OpenAI runs a massive, production-grade product. Yet both score in the C range. The C+ for Anthropic likely reflects a better public narrative around 'Constitutional AI' and more frequent safety white papers. But the C- for OpenAI reflects a wider product attack surface and a more complex ecosystem. Neither score tells you which model resists adversarial prompting better.
Based on my own audit experience, I've found that security outcomes in AI correlate far more with infrastructure implementation than with policy documents. A team that manages its prompt-injection and log monitoring pipelines effectively is safer than a team that publishes 300 pages of safety guidelines and fails to sanitize its inputs. The index ignores this. It creates a false signal that the company with the better PDF is the safer model. That's a critical blind spot for enterprise buyers.
Contrarian: The Military Boogeyman and the 'Safety Tax'
The report also flags the deepening ties between AI companies and the military. This is a classic ethical red flag. But the contrarian angle is that military contracts often demand more rigorous security audits than commercial ones. A defense department contract usually has penetration-testing requirements, data isolation mandates, and adversarial testing protocols that consumer products never face. A company building for the military might have a stronger technical safety post than a pure consumer company. The score punishes the optics, but the actual risk might be lower.
The other blind spot is the 'safety tax.' The narrative suggests that a higher safety score will translate to a business premium. But in the current market, enterprises are not paying a premium for governance. They are paying for latency, accuracy, and price. OpenAI's C score has not stopped it from dominating enterprise adoption. Anthropic's C+ hasn't given it a decisive market share lead. The index has not become a procurement gatekeeper. It is still a niche report for academics and niche news outlets.
Takeaway: Watch the Blind Spots, Not the Grades
Stop treating these grades as a proxy for security. If you are a developer integrating an API, the only relevant data is your own red team tests. The only law that matters is your own stress test. If you are a regulator, the score is useful for understanding a company's willingness to communicate, but it tells you nothing about its ability to contain risks. The next black swan event will not come from a company with a 'C' in a governance report. It will come from a critical overflow bug in a dependency that no one audited. As for the military angle, watch the contracts, not the headlines. The most dangerous code is often the code that compiles without mercy—and without a single report.