Last updated: July 26, 2026 · By Vishal Swami, Founder & Lead AI Reviewer, AISagely
The UK AI Security Institute and the US Center for AI Standards and Innovation jointly tested Kimi K3. They checked Moonshot AI's new open-weight model for offensive cyber skill and published the results on July 23, 2026. Kimi K3 scored well below frontier US models on every measure the two agencies ran. It still showed a real, if limited, ability to write exploits and chain attack steps on its own.
Short answer: UK AISI and CAISI's joint assessment found Kimi K3 scored 32.2% on their aggregate cyber capability measure versus 76.2% for top US models, completed 0 of 41 arbitrary code execution tasks versus a 20-of-41 average for those models, and reached step 17 of a 32-step simulated attack path in 1 of 10 attempts.

In my testing of the actual primary documents behind this story, I read the AISI blog post and the matching NIST notice, not the aggregator summaries quoting them secondhand. The numbers matched across both write-ups. That's reassuring, since a single stat often gets mangled after three or four rewrites. This piece walks through what the assessment measured, what it didn't, and how to read a report like this one without panicking or waving it off.
What you'll need to understand this report
You don't need a security background to follow this. UK AISI is the UK's AI Security Institute, part of the Department for Science, Innovation and Technology. CAISI is the US Center for AI Standards and Innovation, housed at NIST. Both are their governments' frontier-model evaluation bodies, and they've run joint or parallel cyber checks on several frontier releases this year. Kimi K3 itself is Moonshot AI's newest flagship model. It's a roughly 2.8-trillion-parameter mixture-of-experts model released July 16, 2026. Full open weights are due out under a modified MIT license on July 27, 2026. That open-weight date is exactly why the agencies moved fast. Once the weights are public, anyone can fine-tune away the safety filters Moonshot built in.
Step-by-step: how to read this assessment for yourself
1. Start with the primary source, not the headline
Go to the AISI blog post directly, or the matching NIST notice, rather than a summary of a summary. Aggregator sites tend to flatten "scored below frontier US models on preliminary cyber evaluations" into "Kimi K3 can hack anything." Others flatten it the other way, into "Kimi K3 is basically harmless." Neither is what the report says.
2. Separate the aggregate score from the individual benchmark
The 32.2%-versus-76.2% figure is an aggregate across multiple tests, built with an Item Response Theory-style scoring approach. The individual benchmarks tell a more specific story: on ExploitBench, a 41-task exploit-development suite built with Carnegie Mellon University, Kimi K3 scored 32% against 24% for GLM-5.2, a Chinese comparator model, but 0 of 41 on the harder arbitrary code execution subset, where the US comparison average was 20 of 41.
3. Check who the comparison group actually is
"Leading US models" in this report is an average of unnamed frontier systems, not one named product. That matters if you're trying to decide whether a specific tool is safer than Kimi K3 — the report simply doesn't name names on the US side.
4. Read the stated caveats before drawing a conclusion
AISI is explicit that its cyber range, called The Last Ones (TLO), "differs from real-world environments in several ways" — it has no active defenders, no defensive tooling, and no penalty for tripping a security alert. A model doing better on TLO doesn't mean it would get further against a real, monitored network.
5. Note the open-weight release date
Kimi K3's weights land July 27, 2026. After that date, the safeguards AISI tested — which it says "did not prevent it from attempting cyber exploit development" even before open release — become someone else's to modify or remove entirely.
The numbers, side by side
| Measure | Kimi K3 | Leading US models (avg) | GLM-5.2 |
|---|---|---|---|
| Aggregate cyber capability score | 32.2% | 76.2% | not reported |
| ExploitBench (41 exploit-dev tasks) | 32% | not reported | 24% |
| Arbitrary code execution (ACE) subset | 0 of 41 | 20 of 41 | not reported |
| TLO simulated attack path (32 steps) | Step 17 (avg) | Step 28.5 (avg) | not reported |
| TLO full-path completions | 1 of 10 attempts | not reported | not reported |
Every figure above comes straight from the AISI and NIST write-ups; I didn't round anything to make the table tidier.
Example prompts you can copy
If you want to interrogate a report like this yourself instead of taking any summary — mine included — at face value, these prompts work in ChatGPT, Claude, or Gemini once you've pasted in the source text:
- "Read this AI safety evaluation and list only the claims that come with a specific number and a named benchmark. Flag anything stated as a general conclusion without a supporting figure."
- "What comparison group is this report using, and does it name specific competing models or just an average?"
- "What limitations or caveats does the report itself state about its test environment? Quote them directly."
- "Summarize what changes once this model's weights are publicly released, according to the report."
- "Draft a two-paragraph internal memo for a non-technical team explaining what this assessment does and doesn't tell us about risk."
Common mistakes to avoid
The mistake I see most is treating "scored below frontier US models" as "can't do anything dangerous" — the report itself says the opposite: Kimi K3 is "capable of autonomously attacking small, weakly defended, and vulnerable enterprise systems, when directed to do so and given initial network access," even at a 1-in-10 success rate. Second, people quote the 32% ExploitBench figure and the 32.2% aggregate figure as if they're the same number measuring the same thing; they aren't, and conflating them muddies every comparison downstream. Third, ignoring that TLO has no active defenders and treating a lab result as a real-world one. Fourth, skipping the date stamp — this is a preliminary assessment of a model whose full open-weight release hadn't happened yet when the report published, so any claim about what third-party fine-tunes can do with it is still speculation. And fifth, assuming "open-weight" and "unsafe" are synonyms; my own testing of open-weight models for this site, covered in my Kubernetes-moment piece on the open-weight ecosystem, keeps turning up the same nuance — capability and safeguard strength vary a lot model to model, and Kimi K3's cyber numbers here are one data point, not a verdict on open-weight models generally.
Tools that make this easier
If your actual interest is which model to use for legitimate work rather than the safety-research angle, my best AI tool for code roundup and ChatGPT alternatives for coding guide both cover models I've tested hands-on for real coding tasks, cyber capability aside. For a closed-model comparison point, how to use Claude AI is the fairest one I've run against open-weight releases like this one. If budget is the constraint keeping you from testing anything yourself, my free AI tools list is a reasonable starting point. And if this kind of story — a specific number getting stretched past what it actually shows — is a pattern you want to get better at spotting, I've written about the same failure mode more broadly in why AI mania is eviscerating global decision-making; my AI tool ratings hub applies the same check-the-primary-source habit to every tool I score.
My take
The UK AISI & CAISI assessment of Kimi K3's cyber capabilities isn't a verdict on the whole model, and I don't read it as either "Kimi K3 is dangerous" or "Kimi K3 is fine." It's a genuinely below-frontier score on cyber-specific tasks, next to a genuinely real, if low-probability, ability to chain an attack unsupervised when handed initial access. Both things are true at once, and the AISI write-up says so directly if you read past the topline percentage. The useful habit isn't picking a side — it's reading the primary source before repeating either version of the headline.
Frequently Asked Questions
Is Kimi K3 free to use?
Moonshot AI is releasing Kimi K3's weights under a modified MIT license on July 27, 2026, which typically means free self-hosting once you have the hardware; check Moonshot's own release terms before assuming zero restrictions.
How long did the UK AISI / CAISI assessment take to run?
The agencies describe it as a preliminary assessment completed and published within about a week of Kimi K3's July 16, 2026 release — fast by design, since the model's open-weight release followed just days later.
What is the easiest way to check a safety report like this myself?
Open the primary source directly — aisi.gov.uk or nist.gov in this case — and read the caveats section before the topline number. That single habit catches most of the misreadings that spread through secondhand coverage.
Does scoring below US models mean Kimi K3 is safe?
No. The report explicitly states Kimi K3's safeguards did not prevent it from attempting exploit development, and that it can autonomously attack small, weakly defended systems when given initial access, even though its overall scores trail top US models.
Who ran this assessment?
The UK AI Security Institute and the US Center for AI Standards and Innovation (CAISI, housed at NIST) ran it jointly, with ExploitBench built in collaboration with Carnegie Mellon University.