Field note · Responsible use

Write the policy before you trust the filter

The “Write Your AI Policy” one-pager

What happened

On 4 August, Mistral released Shieldstral — a small (3B), open-weights safety classifier with one unusual property: you hand it your policy in plain English at inference time, and it scores content against YOUR rules, not a vendor's fixed list. Performance figures are vendor-reported and unverified, but the design shift is real: safety is becoming something you configure with words, not something you inherit from whoever built the model.

Why it matters to you

Every team using AI at work has a policy problem before it has a tooling problem. "Don't put client data in ChatGPT" is not a policy — it's a hope. The moment tools like this exist, the bottleneck moves to something only you can supply: a written, testable statement of what's acceptable in YOUR workflow. Teams that have one can adopt tools like this in a day. Teams that don't will discover their "policy" was never specific enough to be enforced by anything — machine or human.

How to do it (today, no installation)

Use the one-page exercise below. It's the same artifact we use in class, and it works even if you never touch the tool — because the value is in the writing, not the classifier.

ARTIFACT — "Write Your AI Policy" one-pager

  1. One sentence of scope. "This policy covers [task] done by [role] using [tool]." Narrow beats noble.
  2. Three allowed / three forbidden. Concrete examples from your own work, not categories. "Summarising our published research: allowed. Pasting a client's portfolio statement: forbidden."
  3. Five ambiguous cases. The real test. Write five inputs where reasonable colleagues would disagree (an anonymised client email? a screenshot with a name half-visible?). Decide each one now, in writing.
  4. The threshold rule. Complete: "When unsure, the default is ___ and the person who decides is ___."
  5. The audit line. Where does the record live? A policy nobody can check was never a policy.

Twenty minutes, one page. If you later deploy a policy-adaptive filter, this page becomes its configuration. If you never do, it's still the clearest AI-governance document your team owns.

The lesson

Safety tooling is becoming programmable — which means governance is becoming a writing skill. The teams that can state their rules precisely will get to use the fast tools safely. The rest will keep arguing about hypotheticals.