# The TypeSafe Agent Skill Teaches Agents to Write Jev Code. Letting Jev Steer Codex Takes Your Own Rules

**The TypeSafe agent skill, tagged version 0.5.7 on September 12, 2026, installs in Codex with one command, or two in Claude Code, and teaches agents to build features on Jev.** It does not make your coding agent hand its own decisions to Jev. That second pattern is possible, but you have to write the prompt, the threshold, the failure handling and the log yourself.

A widely shared Codex and Jev demo shows that second pattern. Codex installs the skill, then asks Jev whether its next move should be a web search, a code read or a question for the user, and falls back to human review below 0.8 confidence. It works as a demo. It also blurs what the official skill is for. If you want the basics first, our earlier [explainer on Jev's typed decisions and pricing](https://developer.tenten.co/what-is-jev-a-guide-to-typesafe-ai-s-system-one-model) covers the three question types. This piece covers what the skill changes inside your agent, and what to add before an agent takes orders from a classifier.

#### Install in one command, then check Codex's key and sandbox settings

TypeSafe publishes the TypeSafe agent skill from its GitHub repository under an MIT license. Claude Code installs it as a plugin; Codex and most other agents use the skills CLI:

```bash
# Claude Code
claude plugin marketplace add typesafe-ai/skills
claude plugin install typesafe@typesafe-ai

# Codex and other agents
npx skills add typesafe-ai/skills --skill typesafe-ai
```

Per the [agent skill installation guide](https://docs.typesafe.ai/agent-skill), the `npx` route installs into the current project by default, and `-g` makes it global. Pick one method. Two copies drift apart when only one gets updated.

| Method | Default scope | Explicit invocation | Update |
|---|---|---|---|
| Claude Code plugin | user scope; [`--scope project` changes it](https://code.claude.com/docs/en/discover-plugins) | `/typesafe:typesafe-ai` | `claude plugin marketplace update typesafe-ai`, then `claude plugin update typesafe@typesafe-ai` |
| `npx skills add` for Codex | project `.agents/skills/` | name the TypeSafe skill in the prompt | `npx skills update` |
| Manual copy | wherever you put it | name it in the prompt | replace the folder with the latest GitHub version |

What you install is small. As of October 5, 2026, the repository holds two plugin manifests, a README, two license files and one 149-line `SKILL.md`. There are no scripts, no hooks and no MCP server. Nothing runs in the background; the agent just gains a document it can read.

Calling the API takes a key, read from an environment variable:

```bash
export TYPESAFE_API_KEY="your-key"
```

Codex adds two wrinkles. Per the [Codex configuration reference](https://learn.chatgpt.com/docs/config-file/config-reference), its `shell_environment_policy.ignore_default_excludes` setting defaults to true, so variables with KEY, SECRET or TOKEN in the name are kept before other filters run. Filters you add later can still drop them, and if you hardened the setting to false, `TYPESAFE_API_KEY` disappears too. The default workspace-write sandbox keeps [network access off](https://learn.chatgpt.com/docs/agent-approvals-security), so Codex asks for approval before a command reaches `api.typesafe.ai`. To let calls through, set `sandbox_workspace_write.network_access = true`. To restrict them to TypeSafe, also enable `features.network_proxy` with an allow rule for the domain. An allowlist alone does not turn the network on.

#### The skill loads on its description, so name it anyway

Claude Code and Codex both decide on their own whether a skill applies. [Claude Code's skills documentation](https://code.claude.com/docs/en/skills) says Claude uses skills when relevant, and a skill's description helps Claude decide when to load it. OpenAI's Codex documentation describes the same implicit invocation, which a skill can turn off with `allow_implicit_invocation: false`.

The TypeSafe description targets three situations. A feature needs programmable common sense. You are brainstorming what AI could add to an app. Or an LLM prompt-and-parse step could become a structured decision. The demo claims Codex will reach for the skill on any judgment, classification or routing task from the next turn on. The documentation is narrower. TypeSafe lists "the agent isn't using the skill" as a common issue and recommends `/typesafe:typesafe-ai` in Claude Code or naming the skill elsewhere. Naming it costs nothing and removes the guesswork.

#### The skill tells your agent to build features and keep logic in code

The first instruction in the TypeSafe agent skill's `SKILL.md` is to read the live docs, starting with the `llms.txt` index, before writing any integration. The agent then works backward from the behavior the user wants to the judgments that behavior needs.

The division of labor is explicit. Known rules, calculations, exact lookups and execution stay in code. Jev supplies semantic judgment through Choice, Noul or Score. Independent questions about the same state go into one request. The agent tests representative cases and separates missing evidence, model errors, code bugs and service failures when something breaks.

On thresholds, the file is blunt. Choice and Score confidence "summarizes distribution concentration, not overall workflow correctness or permission to act." Cookbook thresholds and demo results are examples to evaluate. TypeSafe's page on coding agents goes further: Jev is not a replacement for the LLM behind Claude Code, Cursor or similar coding tools, and no setting turns a coding agent into a Jev-powered one.

So the intended user is a developer adding ticket routing, reranking or moderation to a product, with the agent writing the integration.

![Two patterns for the TypeSafe agent skill: the agent writes code that calls Jev, or hands its own next step to Jev behind a threshold gate](https://s4.tenten.co/learning/content/images/2026/10/linkedin-infographic-1-8.png)

#### The viral demo hands the agent's own next step to Jev

The widely shared Codex and Jev demo sends Codex's current state to a Jev Choice question. The user asks why a stock rose today, and the known facts list is empty. There are five candidate actions: answer directly, search the web, inspect local code, ask the user, or escalate to human review. Codex acts only at confidence of 0.8 or higher. Any API failure stops the run.

The demo's author reported `jev-1.13.0`, 495 input tokens, 63 output tokens and a 0.99 probability on `web_search`, with confidence also shown as 0.99. Codex then searched. Those figures are the author's own record; Tenten did not rerun the demo and cannot verify them. Checked against the documentation, four things stand out.

The pasted JSON is the agent's summary, not the API response. The [HTTP API reference](https://docs.typesafe.ai/api) nests each result under `answers`, keyed by your question ID, with `choice`, `probabilities` and `confidence` inside. There is no `finalAction` field. If you want an audit trail, have the agent save the raw body.

The 0.8 bar needs translating. TypeSafe's [Choice confidence formula](https://docs.typesafe.ai/confidence) is (top probability − 1/n) ÷ (1 − 1/n). With five options, 0.8 confidence means a top probability of 0.84. A 0.99 probability works out to 0.9875 confidence, which rounds to the reported 0.99.

The case is easy. An empty fact list and a question about today's market leave almost any router choosing search. The run proves the wiring works. It says nothing about whether 0.8 is the right threshold.

The model cost really is small. Jev 1.13 is priced at [$0.042 per million input tokens](https://docs.typesafe.ai/models), and output is free. At the demo's 495 tokens, one decision costs about $0.0000208, and 10,000 decisions cost about $0.21. That's our arithmetic, and it covers Jev only. The tokens Codex spends writing the request and reading the answer are extra, and we have not measured whether the pattern saves Codex quota overall.

| | Official skill pattern | Demo pattern |
|---|---|---|
| What Jev judges | Data your application handles | The coding agent's own next step |
| Where the threshold lives | A reviewed constant in code | One line in a prompt |
| Who handles failures | Your application code | The agent, if it follows the prompt |
| Cost per decision | Jev input tokens | Jev input tokens plus the agent's own tokens |

#### Four guards before Codex takes orders from Jev

Letting Codex hand its next step to Jev is worth trying when the options are a few low-risk moves plus a way to ask a person. Here is a prompt rewritten against the documented API fields. It is an example; Tenten has not run it or called the paid API with it.

```text
Use the TypeSafe skill.

Jev decides this step. Call POST https://api.typesafe.ai/v1/systemone for real; do not answer from your own analysis.
Use model jev-1.13.0. Read the key from TYPESAFE_API_KEY and never print it.

state:
{
  "userRequest": "Explain why NVIDIA stock rose today",
  "knownFacts": [],
  "availableTools": ["answer_directly", "web_search", "inspect_code", "ask_user"]
}

Question (Choice): To complete the request reliably, which action should come next?
Options:
- answer_directly: stable information is already available
- web_search: the answer depends on current outside information
- inspect_code: local project files must be checked
- ask_user: information required to finish the task is missing
- human_review: the risk is high or no option fits

Rules:
- Act on Jev's choice only when confidence is 0.8 or higher; otherwise switch to human_review.
- On 401 or 422: stop and report. Do not retry.
- On 429 or 529: retry with exponential backoff, at most 3 times, then stop and report.
- Never substitute your own judgment after a failure.
- Append the raw model, usage and answers to logs/jev-decisions.jsonl before acting.
```

![With five options, 0.8 Jev confidence equals a 0.84 top probability, plus four guards: pin jev-1.13.0, set thresholds from labeled cases, stop or back off on API errors, log each decision](https://s4.tenten.co/learning/content/images/2026/10/linkedin-infographic-2-8.png)

##### Pin the model version and name the skill

When Codex delegates a step to Jev, pin the Jev model version: TypeSafe's models page says the `jev-latest` alias moves with each release, so answers can change under you. If you tuned a threshold, keep `jev-1.13.0` until you retune. Open the prompt by naming the TypeSafe skill and telling Codex to call the API for real, because the skill may not load on its own.

##### Set thresholds from your own labeled cases

A Jev confidence threshold for agent steps should come from your own labeled cases. Pull past tasks, mark the right next step by hand, and compare error and escalation rates at several cutoffs. TypeSafe's guidance scales thresholds with risk: a read-only search can clear a lower bar than a file write or an outbound message. If you only need the most likely of several harmless options, the docs say to take the top choice. Deciding when to escalate to a person is what needs a threshold.

##### Fail closed on every API error

A Codex-to-Jev decision step should fail closed: TypeSafe's API returns 401 for a bad key, 422 for a malformed request, 429 for rate limits and 529 when TypeSafe is overloaded. The first two won't fix themselves on retry. The last two can back off and try again. Current limits are 80 requests and 100,000 tokens per second, and TypeSafe says those numbers are adjusting dynamically.

##### Log each decision and keep secrets out of state

Every Jev decision an agent acts on needs a log entry: model version, state, options, full probabilities, confidence, threshold and the action taken. Keep keys, personal data and proprietary code out of `state`, because it leaves your machine. TypeSafe says Jev isn't trained on customer requests, and zero data retention is an enterprise option.

None of this replaces the agent's own sandbox. If Jev picks `web_search`, Codex's approval rules still apply. It's the same split we used for a [Grok Bot decision layer](https://developer.tenten.co/grok-bot-jev-decision-layer): the model picks, code checks permission.

#### Where the demo's claims and the documentation part ways

Checked against TypeSafe, OpenAI and Anthropic documentation, one of the demo's five claims matches, one is half right, and three need correction or can't be verified. The speed and price multiples compare Jev with frontier LLMs.

| Demo claim | What the maker's sources say | Verdict |
|---|---|---|
| The skill triggers automatically from the next turn | Loading depends on the description; TypeSafe recommends naming the skill | Possible, not guaranteed |
| 40x to 200x faster, 70ms to 500ms | The launch post gives the same figures | Matches, for System One shaped queries |
| 40x to 400x cheaper | No such figure found; the homepage says 444.6x cheaper and 238x lower input price than Claude Fable 5.1 | The 444.6x comes from TypeSafe's own workflow evals |
| The response has `choice` and `finalAction` | Answers sit under `answers`; no `finalAction` field exists | The demo showed an agent summary |
| The open-source clone Kev was written with Devin | The Kev repository exists; its README never mentions Devin | Authorship claim unverified |

Jev's speed figures come from TypeSafe's [September 15, 2026 launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev). The same post says the homepage's 193.6x and 444.6x numbers sit at the high end of real-world gains, and that TypeSafe's own capabilities team built the test workflows.

#### FAQ

##### How do you install the TypeSafe agent skill in Codex?

The TypeSafe agent skill installs in Codex with `npx skills add typesafe-ai/skills --skill typesafe-ai`, choosing Codex when prompted. The default target is the project's `.agents/skills/` folder; `-g` installs it for every project. If Codex still ignores the skill, TypeSafe suggests checking the install target and restarting the agent.

##### Does the TypeSafe agent skill replace the model behind Codex?

No, the TypeSafe agent skill leaves Codex's model unchanged. TypeSafe's documentation says Jev does not generate text or code, and no setting turns a coding agent into a Jev-powered one. The skill only teaches the agent the API, the question types and the design rules.

##### Do you need a TypeSafe API key to install the skill?

No, installing the TypeSafe agent skill needs no key, and the repository contains no executable code. A key in `TYPESAFE_API_KEY` matters only when the agent calls Jev or runs test queries. In a web app, the key belongs on the server.

##### Can Kev serve the same decision requests as Jev?

Kev can accept TypeSafe's System One request format, but thresholds tuned on Jev may not carry over, so retest them. Kev is an Apache-2.0 project on the jaredpalmer GitHub account, built on Qwen3.5 and Qwen3.8 base models. Its README says TypeSafe's Python SDK works against a Kev server unchanged.

#### Sources

- [TypeSafe AI: typesafe-ai/skills repository](https://github.com/typesafe-ai/skills)
- [TypeSafe AI: Agent skill installation and troubleshooting](https://docs.typesafe.ai/agent-skill)
- [TypeSafe AI: Jev with coding agents](https://docs.typesafe.ai/introduction/coding-agents)
- [TypeSafe AI: HTTP API reference](https://docs.typesafe.ai/api)
- [TypeSafe AI: Confidence](https://docs.typesafe.ai/confidence)
- [TypeSafe AI: Models, pricing and rate limits](https://docs.typesafe.ai/models)
- [TypeSafe AI: Introducing System One models and Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
- [TypeSafe AI: Homepage speed and price claims](https://typesafe.ai)
- [Anthropic: Claude Code skills](https://code.claude.com/docs/en/skills)
- [Anthropic: Discover and install Claude Code plugins](https://code.claude.com/docs/en/discover-plugins)
- [OpenAI: Build skills for Codex](https://learn.chatgpt.com/docs/build-skills)
- [OpenAI: Agent approvals and security](https://learn.chatgpt.com/docs/agent-approvals-security)
- [OpenAI: Codex configuration reference](https://learn.chatgpt.com/docs/config-file/config-reference)
- [jaredpalmer/kev repository](https://github.com/jaredpalmer/kev)

#### Author Insight

A coding agent's next step is a worse place for an unattended confidence threshold than a support queue. A misrouted ticket lands in the wrong inbox; a misrouted agent step can edit code or send something outside. TypeSafe's own docs say confidence is not permission to act. Until you have labeled history for your agent's decisions, Jev should only gate read-only moves like search and file reads, and anything that writes should still wait for a person.

