What Jev (TypeSafe) Means for Security Operations

What Jev (TypeSafe) Means for Security Operations

Ajmal Kohgadai
Ajmal Kohgadai
September 29, 2026

Jev is an AI model that answers a question by choosing from a fixed set of answers and reporting how confident it is, without writing any text. TypeSafe AI released it on September 15, 2026. The company named it after William Stanley Jevons. In 1865, Jevons argued that more efficient steam engines would make Britain burn more coal, because cheaper power gets put to more uses.

TypeSafe chose the name as a forecast: "We expect machine intelligence to follow a similar path to coal." The same logic applies to security operations. TypeSafe says Jev answers in 70 to 500 milliseconds and charges $0.042 per million input tokens, with output free. At that price, a team can afford to score a much larger share of the process events, emails, and logins it collects, instead of only the ones that fire an alert. If the Jevons pattern holds, security teams will make far more automated judgments than they do today. Each judgment that comes back suspicious will still need an analyst, or an AI SOC analyst, to find out what happened.

A SOC team looking at Jev needs to know four things: what TypeSafe shipped, which security decisions suit a System One model, what a Jev answer cannot tell an analyst, and what to test first.

{{ebook-cta}}

What TypeSafe released with Jev, and what is still unverified

Jev is the first of what TypeSafe calls System One models. The label comes from Daniel Kahneman's book Thinking, Fast and Slow, which separates fast, intuitive System 1 thinking from slow, deliberate System 2 reasoning. TypeSafe built Jev for the fast side: decisions that software makes inside a workflow and acts on directly.

A developer sends Jev a block of text or JSON, which TypeSafe calls the state, along with one or more questions. There are three question types. A Choice picks one option from a list of up to 255. A Score rates the input on an ordered scale. A Noul returns the probability that a statement is true. Every answer comes back with probabilities, and Choice and Score answers also carry a confidence value. Jev produces all of its answers in a single query, so there is no text to parse and no answer outside the options the developer defined.

The performance numbers come from TypeSafe. The company reports Jev at 40 to 200 times faster than frontier models on comparable tasks, and at 193.6 times faster and 444.6 times cheaper on one set of workflow evaluations. TypeSafe also says its own model team wrote those workflows, that "some bias could exist," and that the figures sit "on the higher end of real world gains." TrueFoundry's review listed those figures as "self-reported and not yet independently reproduced."

The claim that needs the most care is that Jev "can't hallucinate." Because Jev never writes free text, it cannot invent a tool name, a hostname, or a citation, and every answer fits the format the developer declared. However, it can still choose the wrong option with high confidence. As TrueFoundry put it, "what's been eliminated is the malformed answer, not the mistaken judgment." In a SOC, the mistaken judgment is the costly failure: an alert marked benign that was malicious.

TypeSafe's documentation is open about where the current model is weak. Jev "answers the question you wrote, not the one you meant," according to the docs. It does not count reliably. It reads dates as text rather than as ordered values, and its accuracy drops as the input fills with content unrelated to the decision. It also accepts text only, with no images.

Which security decisions a System One model handles well

Many of the decisions inside a security workflow are bounded, meaning the answer is one item from a known list. These are the decisions Jev was built for, and they come up at every stage of alert triage:

  • Routing: sending an alert to the investigation plan, queue, or team that matches its type.
  • Tagging: labeling activity with a MITRE ATT&CK tactic or an alert category so it can be grouped and reported. Many detection rules already carry ATT&CK tags, as Sigma rules do in their tags field, so this helps most with raw events and alerts that arrive untagged. A technique can also fall under more than one tactic. Valid Accounts falls under four. A Choice returns one option, so tagging several tactics takes a separate Noul question for each one.
  • Ranking: ordering candidates by relevance, such as the knowledge-base articles, related alerts, or asset owners worth pulling into an investigation.
  • First-pass scoring: rating high-volume telemetry, such as the process events behind EDR alerts, before any of it becomes an alert. Neither an analyst nor Jev can usually judge a process event on its own. Microsoft notes that malware commonly abuses Office apps to start child processes, and that some legitimate business apps start PowerShell for benign reasons. The state needs the parent process chain and the host's role, gathered before the call.
  • Guardrails: screening the prompts, tool outputs, and responses of an LLM-based agent for prompt injection before another model acts on them. TypeSafe lists this use case directly and publishes a guardrails recipe for it.

At a small fraction of a cent and well under a second per decision, a team can score events that were never worth a call to a frontier model. Price is not the only ceiling. TypeSafe currently lists a default rate limit of 1,200 requests a minute, about 1.7 million a day, with higher limits on custom and enterprise plans. False positives are the bigger problem. If a team scores 10 million events a day and 0.1% come back as false positives, that is 10,000 flags a day for someone to work.

For security work, the confidence value is the most useful part of the output. TypeSafe documents a pattern it calls confidence-gated routing: accept the answers the model is confident about, and send the rest to a slower model or a person. A SOC team using that design lets Jev handle the clear-cut cases and keeps the ambiguous ones for an analyst or a model that can reason. It only works if Jev's confidence scores match how often its answers are right. TypeSafe claims its training method produces calibrated confidence, and each team should confirm that on its own data.

What a Jev answer cannot tell an analyst

Jev returns a label and a probability for a question you define in advance; it does not decide which questions an alert needs, gather the evidence, or show its reasoning. When an analyst reviews an automated verdict, the first thing they look for is what was checked: which queries ran, what came back, and why that evidence supports the call. A score of 0.91 that a process is suspicious gives the reviewer nothing to verify. For that reason, security teams evaluating AI for the SOC ask to see the queries behind each determination, and explainability is a standard evaluation question.

Jev also answers only the questions it is given. In an investigation, the analyst picks the next question based on the last answer. Take an Okta alert for a login from a new country. The analyst checks whether MFA was satisfied, then what the session did after login, then whether the user can confirm the trip. Each step depends on what the previous one found. Jev can answer any of those questions quickly once something asks it. Choosing the questions, running the queries that produce the evidence, and connecting the answers into a determination is investigation work. That work has to come from an analyst or from an agent built to plan investigations.

Two limits in TypeSafe's documentation affect alert handling directly. First, many security decisions depend on counts and time. Examples include the travel speed implied by two logins in an impossible travel alert, the number of failed attempts in a window, and the minutes between an email's delivery and the first click. TypeSafe says to compute values like these in code and send Jev the result. Second, Jev accepts text only, so a phishing investigation that turns on a screenshot, a QR code, or an image attachment needs another tool for that part.

What to check before Jev touches your alert queue

A team putting Jev in an alert path should hold it to the same standard as a proof of value for any AI SOC tool. That means testing on its own data, against known answers, with the failure cases in view. These checks apply before any production use:

  • Include malicious activity in the test set. A sandbox or a quiet tenant can hold weeks of events with nothing confirmed malicious in them. Results on that data show how Jev handles routine activity and say nothing about the alerts that matter most.
  • Grade against confirmed outcomes. Scoring Jev against a frontier model's labels measures how often the two models agree. Where they disagree, either one can be wrong, so each disagreement needs an analyst to decide which answer is right.
  • Check calibration by confidence band. Group answers by confidence and measure accuracy in each group. If answers at 0.9 confidence are right about nine times in ten on your data, the confidence value can be used to route work. If they are not, the gate needs a higher threshold or a different design. Even when the scores are calibrated, answers at 0.9 are wrong one time in ten, which is too many for alerts closed without an analyst. TypeSafe's guide to confidence-gated routing says to set each threshold by the cost of acting on a wrong answer. Also count the confirmed-malicious items that land in the band you plan to accept automatically, since those are the misses that cost the most.
  • Send named fields. TypeSafe's documentation says accuracy falls as unrelated content grows, and cost rises with every input token. Pass the fields a decision depends on, such as the command line, parent process, and user, rather than the full alert payload.
  • Treat attacker-written text as hostile input. The attacker writes the command lines, email subjects, file names, and script comments. TypeSafe's documentation says Jev does not treat the state as hostile by default, and that text written to steer the model "can move the answer." Test whether a comment naming a credential-dumping tool raises a benign command's score, and whether a forged "approved by security" note lowers a malicious one. Where either one works, clean those fields before scoring. Remove comments and decode encoded commands, but keep the command line itself, since the score depends on it.
  • Pin the model version. TypeSafe's jev-latest alias moves to each new release, so answers can change without any change to your code. Thresholds tuned on one version need retesting on the next, and each result should be logged with the versioned model name.
  • Read the data terms before sending production data. Alerts carry usernames, hostnames, and email content. TypeSafe commits not to train on user data and offers zero data retention to enterprise customers through its sales team. Confirm which terms apply to your account, and put a data processing agreement in place first.

What cheaper decisions mean for the SOC

If Jevons was right about coal and TypeSafe is right about machine intelligence, the cost of a single security judgment will keep falling and teams will make more of them. When scoring costs almost nothing, a SOC's capacity depends on how many flagged items it can investigate and document. A cheaper classifier can clear the easy items from the queue, but an analyst or an agent built to investigate still has to work the rest.

Prophet AI SOC Analyst was built for that part of the work. It plans an investigation for every alert and queries the SIEM, EDR, identity, cloud, and email tools a team already runs. It records every query and piece of evidence behind its determination, and it returns Inconclusive when the evidence does not support a call.

Table of contents
Add as Google Preferred Sources

Insights

Definitive Guide to AI SOC Agents

This guide breaks down how AI SOC agents work and how to build an agile security operation around agentic AI

Download eBook
Ajmal Kohgadai

Ajmal Kohgadai

As the Director of Product Marketing at Prophet Security, Ajmal drives marketing and growth strategies and helps security professionals see how AI is transforming security operations.