LLM Unlearning, Security Hardening

Enterprises are Shifting to Open-Weights. Now It's Time to Harden Them

By
Yael Kishon
July 22, 2026

According to the State of Open Source AI 2026 report by Mozilla, open-weight models have moved from experimental to production: 

  • Open-weight models now account for a majority of production inference tokens on OpenRouter 
  • The top three blockers to adopting open-weight models: infrastructure costs, security, privacy, or compliance concerns, and ongoing maintenance
  • Open model performance keeps improving — the gap with top proprietary systems (ChatGPT and Claude) is now just 3.3%
  • China and East Asia now lead the world in open source AI adoption

Leading enterprises are now turning to Chinese open-weight models

This isn't theoretical. Some of the most operationally demanding companies in the world are betting on open weights:

  • Shopify fine-tuned Qwen3-32B to power a tool-calling agent inside Sidekick, its AI commerce assistant. The reasoning is gaining competitive differentiation: closed models commoditize any capability accessible via API key, so lasting advantage has to come from owning the proprietary data, training recipe, and infrastructure behind the model. The result was a model 2.2x faster, 68% cheaper, and more accurate than closed alternatives.
  • Harvey  the legal AI company, built its own cloud agent infrastructure specifically to run open-weight models — driven by client confidentiality requirements, zero-data-retention mandates, and cost. On Harvey's Legal Agent Benchmark (LAB), GLM 5.1 is the strongest open-weight performer, outperforming Opus 4.7 and GPT-5.5 — proof that open models are closing the gap on frontier legal work, not just cheap fallback options.

Both are proof of the same underlying stack. The diagram below shows where each piece sits: agents on top, an agentic harness orchestrating them, the open-weight models underneath doing the actual reasoning, and cloud infrastructure carrying it all. Hirundo operates at that model layer specifically — not inside the harness, not down in the infrastructure, but on the weights themselves, hardening whatever model Shopify, Harvey, or anyone else chooses to run.

Hirundo operates at the model level — making open-weight models safer


The tradeoff nobody's pricing in: security


Security is already the second-biggest concern holding open-model adoption back. Security, privacy, or compliance concerns ranked second among barriers at 26% - just behind infrastructure cost (27%), and ahead of maintenance (24%) and deployment complexity (23%). In South Asia, that concern spikes to 39%. Enterprises aren't imagining this risk; they name it as one of the top two reasons open-weight adoption never makes it to production.

Meanwhile, Chinese open-weight models are being adopted at scale. Chinese models grew from under 2% of weekly token traffic in late 2024 to more than 45% by April 2026. Alibaba's Qwen alone racked up 942 million downloads by March 2026.

That's why Hirundo chose Qwen as its unlearning target: it's one of the models being deployed at scale, and it's a model where security and compliance concerns are most acute - geopolitical sensitivity, unclear data origins, and the black-box nature of a model users can't fully inspect.

Unlearning as the hardening layer

This is where Hirundo's unlearning approach comes in: we remove the underlying knowledge from the model's weights directly.

Case study: Qwen3.5-27B

Hirundo ran a dedicated security hardening pass on QWEN3.5-27B, targeting prompt injection vulnerabilities at the weight level. On the Purple Llama CyberSecEval benchmark, attack success rate dropped from 13.9% to 2.79% — a 80% reduction.

Utility Preserved

Usually, removing something from a model risks damaging something else along with it. However, Hirundo preserved performance across the full benchmark suite — in fact, the hardening produced a net improvement, with AIME 2025, MMLU-Pro and SciCode all gaining ground and only negligible drops on GPQA, IFBench and LiveCodeBench.

If you’re deploying QWEN3.5-27B or any other model in a production environment and want to understand your exposure — or eliminate it — get in touch. We’re happy to run an assessment.

Yael Kishon
Product Manager

Ready to forget?

Start removing unwanted data with a few clicks