codingBy HowDoIUseAI Team

4 ways to use Jev that actually make your AI coding agent better

Jev isn't a chatbot - it's a tiny decision model that guards, tests, and routes your coding agent. Here's how to actually use it.

Your coding agent just asked to delete a folder, read your .env file, and run a shell command with sudo in it — all in the span of ten seconds. Do you want to pause and review every single one of those calls with a full-sized LLM? That would be slow and expensive. Do you want to just let them through? Absolutely not.

This is the exact problem Jev was built to solve, and it's why it's quietly become one of the more practical additions to serious AI coding setups this year. Not because it's flashy — it doesn't write code, doesn't chat, doesn't generate anything at all — but because it does one very narrow thing extremely fast and extremely cheap: it makes decisions.

If you've only seen Jev in viral demo clips, you've probably missed what makes it actually useful. This guide covers four specific, practical ways to wire Jev into your coding workflow so your agent runs safer, cheaper, and with way fewer wasted tokens.

What is Jev, actually?

Jev is TypeSafe AI's first "System One" model, and it's fundamentally different from something like Claude or GPT. Instead of generating text one token at a time, it takes unstructured state in and returns typed decisions with calibrated probabilities, all sampled in parallel. It's positioned as a fast, cheap decision function for automation code, not a chat model.

Every request to Jev follows the same shape: you send it some state (text, JSON, whatever describes the current situation) along with a set of typed questions, and it answers each one. The API exposes a small set of question types: Choice for categorical classification, Score for numeric or rubric scoring, and Noul for yes/no probabilities.

The speed and cost difference compared to a frontier LLM is the whole point. TypeSafe's Jev is a System One model that returns typed decisions instead of text, running 40-200x faster and 40-400x cheaper than frontier LLMs like GPT-6. And because the output space is bounded ahead of time, it cannot hallucinate or produce type errors, because valid outputs are defined in the schema in advance.

That tradeoff — narrower capability for massive speed and reliability gains — is exactly what makes it a good fit for the repetitive, high-volume judgment calls your coding agent makes constantly. Here are the four places it earns its keep.

How do you use Jev as a security guard for your coding agent?

This is the single most useful integration, and it plugs directly into the hook system most agentic coding tools now support. One of the most practical uses of Jev is as a guardrail inside the hook system that most AI coding agents now support, a pattern popularized by Claude Code and since adopted broadly. Hooks let you attach automations to events in an agent's lifecycle, and the most useful one for security is "pre-tool-use": the moment right before an agent executes an action like reading a file or running a shell command.

Claude Code's hooks documentation lays out exactly how this works under the hood. The PreToolUse event fires before any tool call executes and can return a structured decision to allow, deny, or ask for confirmation — with "deny", Claude Code cancels the tool call and feeds permissionDecisionReason back to Claude. That's the exact slot Jev drops into: instead of writing brittle regex rules for "dangerous commands," you hand the proposed action to Jev and ask a simple Noul question like "does this command delete files outside the project directory, or exfiltrate data?"

The economics make this a no-brainer at scale. Running every single tool call through a frontier model would be absurd, but Jev changes the math — the endpoint is POST https://api.typesafe.ai/v1/systemone, and the early-access model route is referenced as jev-latest, and at roughly a tenth of a cent per analysis, you can screen thousands of calls a day without blinking.

Real projects are already built around exactly this pattern. One open-source hook, jev-axi, is a PreToolUse gate for Claude Code and Codex that has Jev score each shell command for destructiveness, exfiltration, remote code execution, and security weakening, deciding routine commands locally so nothing is sent for them, and scoring 44/44 on the 44 labeled tool calls in its repository. Another, jev-secret-guard, blocks known key formats locally and sends unknown high-entropy strings to Jev only in masked form for a Noul on whether they are real credentials, blocking at 0.80 and asking the human from 0.30 or whenever Jev is unavailable.

How to set this up:

  1. Read the hooks guide to understand the PreToolUse input/output contract.
  2. Get a Jev API key from the TypeSafe quickstart.
  3. Write a small hook script that extracts tool_input from the hook payload, sends it to POST /v1/systemone with a Noul question like "is this command destructive or does it read secrets," and returns permissionDecision: "deny" when the probability crosses your threshold.
  4. Start in shadow mode — log what Jev would block without actually blocking — before flipping it to active. The community plugin jev-tools ships exactly this pattern by default.

Can Jev catch bugs your unit tests miss?

Unit tests are great at catching regressions in logic you already anticipated. They're terrible at catching "the physics feel wrong" or "the enemy got stuck in a wall" — the kind of bug you only notice by actually playing the thing.

This is where Jev's speed becomes a feature rather than just a cost saver. The best pattern here is to sandwich Jev between two LLM calls: have a capable model build the test harness (the scripted playthrough, the scenario setup), let Jev make the fast, repeated judgment calls while that scenario plays out frame by frame or step by step, and then hand any flagged moments back to the LLM to diagnose and fix.

Because Jev can answer dozens of structured questions per second rather than needing a full generation cycle for each one, it can plausibly "watch" a playtest the way a QA tester would — asking "did the player get stuck," "does this look like a collision glitch," "is this score mathematically possible" — far faster and cheaper than routing every frame through a reasoning model. Its core strength, fast and cheap decision-making based on a defined state and options, applies to any domain needing real-time or high-volume decisions, such as game AI, browser automation, and content moderation.

How to set this up:

  • Have your LLM write a lightweight harness that steps through gameplay states (position, score, inventory, screen state) at a fixed interval.
  • At each step, send the state to Jev with a handful of Noul/Score questions tuned to your game's known failure modes.
  • Collect flagged timestamps and states, then pass only those back to your coding agent to investigate and patch — instead of re-running the whole playtest through an expensive model.

How do you use Jev to drive browser testing?

Browser automation has always had an awkward middle step: something needs to look at the page and decide what to click next. Doing that with a full LLM call on every single interaction is slow and burns through tokens fast, especially on multi-step flows.

Jev fits neatly into that decision point. The browser-nav skill in the jev-tools plugin shows the pattern clearly — Jev picks each next click from the page's interactive elements; Claude executes and judges pass/fail. Jev isn't writing the test or interpreting the result — it's just making the cheap, repeated "which element matters here" call, while the LLM retains responsibility for the parts that actually require reasoning: building the test plan upfront and judging success or failure at the end.

That division of labor echoes the general advice from people who've built production Jev integrations: a useful Jev integration leaves open-ended writing and reasoning with the model that is good at those jobs, then uses Jev where a bounded decision must control the next branch. Browser testing is a textbook case — hundreds of "which element, which action" micro-decisions per test run, each one perfectly suited to a Choice question instead of a full generation.

How to set this up:

  • At each page state, extract the list of interactive elements (buttons, inputs, links) as your "state."
  • Send that list to Jev as a Choice question — "which element should be clicked next to accomplish [goal]."
  • Let your agent execute the chosen action and only call back to the LLM when Jev's confidence is low or the test reaches a checkpoint that needs real judgment.

How does Jev decide which workflow your request even needs?

The fourth use case is less about guarding actions and more about routing them. Every coding agent eventually runs into requests that are genuinely ambiguous in scope — is this a quick one-line fix, or does it need a full feature plan with multiple files touched? Deciding that manually, every single time, doesn't scale once you're running dozens of agent sessions a day.

Jev is built specifically for this kind of upfront classification. One documented pattern routes every prompt through a Score question before it ever reaches the main model: questions defined with a "difficulty" score rating how much reasoning a request needs, from a lookup or single-file edit up to architecture or debugging across systems, plus a boolean for whether it needs documentation or data outside the repo.

That score then determines everything downstream — whether the request goes to a cheap, fast model or escalates to a frontier one, and whether a "feature" skill gets loaded or a quick "patch" path handles it instead. The community has built several variations on this exact router. One Claude Code plugin asks Jev a Choice over effort levels — low, medium, high, max, unclear — plus a Noul on whether a hands-off request has a fuzzy spec, showing a switch tip before Claude starts only at 0.7 confidence or above, with 95% of tips pointing to the right level on a three-rater held-out set.

LiteLLM has built this directly into its proxy layer, too — configuring a JEV Auto Router with classifier_type set to jev, which uses one System One Choice question for the configured tiers, then dispatches to the selected completion model.

How to set this up:

  1. Define your Score or Choice question around the dimension that actually matters for your workflow — reasoning depth, risk level, or whether a skill/feature needs to be loaded.
  2. Fire that question at the start of every session via a UserPromptSubmit hook, before the main agent processes the request.
  3. Branch your workflow based on the returned confidence — route high-confidence, low-complexity requests to a cheap path, and only escalate to a frontier model or a planning step when Jev isn't sure.

Where do you actually get started?

Start with the TypeSafe AI quickstart, which walks through getting an API key and making your first request to the /v1/systemone endpoint with real sample code. From there:

  • The models reference covers pricing, rate limits, and context window details for jev-latest.
  • Claude Code's hooks guide is essential reading if you're wiring Jev into PreToolUse or UserPromptSubmit events.
  • The awesome-jev-tools GitHub list rounds up community plugins for security gating, model routing, and browser navigation — most of the patterns above already exist as installable hooks rather than something you need to build from scratch.
  • If you just want to see the three question types in action before writing any code, the live Jev playground lets you test Choice, Score, and Noul questions with no signup required.

None of these four use cases require Jev to be smart in the way an LLM is smart. They require it to be fast, cheap, and boring — the same kind of boring that makes a circuit breaker boring, right up until the moment it saves you from a very bad afternoon. Wire it into the decision points your agent hits hundreds of times a day, and let your expensive model stay focused on the parts of the job that actually need it to think.