How we use Jev to run a GTM autopilot
Most of the work in marketing automation is not writing. It is deciding: is this thread worth answering, what does this reply mean, how warm is this lead, is this draft safe to send. Here is how AutoTraction hands those decisions to Jev, TypeSafe's decision model, with the actual request shapes and thresholds.
A go-to-market autopilot spends very little of its time writing. It spends most of its time deciding. Is this Reddit thread one we should answer? Does this subreddit even allow it? Is the person who replied to our cold email interested, annoyed, or on holiday? Which of the four hooks we drafted should go out this morning? Is this draft safe to post under a founder's name? Only after all of that does anything get written, and the writing is the easy part.
Until this month, AutoTraction made every one of those decisions the way everyone does: send a prompt to a large language model, ask it to answer in JSON, parse the JSON, and hope the "confidence: 0.8" it wrote back means something. On 15 September 2026 TypeSafe AI opened early access to Jev, a model built for exactly this class of problem. This article is about where it fits in a marketing autopilot, with the real request shapes, and what stays with the LLM.
What is Jev?
Jev is a decision model. You give it some state (a thread, a reply, a lead, a draft) and a list of typed questions, and it returns typed answers with probabilities instead of text. TypeSafe calls this a System One model, after the fast, intuitive mode of thinking in Kahneman's framing, as opposed to the slow, reasoning mode that chat models imitate. It does not generate tokens one at a time; it produces the answer in a single pass, which is why TypeSafe claims it is around a hundred times faster and cheaper than a typical LLM call. Input is priced at $0.042 per million tokens and output is free.
There are exactly three kinds of question you can ask it:
| Primitive | What it returns | Use it for |
|---|---|---|
| Noul | A probability, 0 to 1, that a yes/no condition holds | "Is this thread worth answering?" |
| Choice | One option from a set you define, plus the probability of every option and a confidence number | "What does this reply mean?" |
| Score | A position on ordered levels you describe, plus probabilities and confidence | "How urgent is this person's problem?" |
That is the whole surface. It runs on TypeSafe's API directly, and it is available through Vercel AI Gateway and Cloudflare Workers AI, which is where we call it from. Independent questions over the same state go in one request and are answered in parallel.
Where does a GTM autopilot actually make decisions?
We went through every place AutoTraction's cron jobs branch on the output of a model. The list below is that inventory. Every row used to be a prompt-and-parse call to a chat model. Every row is now a Jev question, and the chat model only gets involved in the last column when there is something to write.
| Decision | Primitive | What the code does with the answer |
|---|---|---|
| Is this Reddit or X thread a real request for what we sell? | Noul | Above threshold, queue a reply for drafting |
| Does this subreddit's rule set allow a product mention here? | Noul | Decides whether the draft may name the product or must just help |
| Which angle fits this thread? | Choice | Picks the brief the writer model receives |
| What does this cold-email reply mean? | Choice | Routes to follow-up, founder inbox, or unsubscribe |
| How warm is this lead, on four dimensions? | Score, four times | Weighted in code into a rank for the outreach queue |
| Which of these hook variants should post today? | Choice | Selects one; the rest stay in the queue |
| Is this draft safe to send under the founder's name? | Noul, several | Any failing check blocks the post and asks for approval |
| Does this mention need a human now? | Noul | Escalates to the founder's phone |
Example: deciding whether to answer a Reddit thread
The four-minute routine starts with a scan of new threads where someone asked for what you built. The autopilot does this scan for the founder, and the hard part is not finding the threads. It is that nine out of ten matches for a keyword are not requests at all. Someone mentioning "landing page tool" in a rant is not asking for one.
Here is the request, trimmed. The state carries the thread, the product, and the subreddit's rules on promotion, which we hold as data (our subreddit rules table is the source). Three independent questions ride in one call.
POST https://api.typesafe.ai/v1/systemone
{
"model": "jev-latest",
"state": {
"thread": {
"subreddit": "r/SaaS",
"title": "How are solo founders handling marketing without hiring?",
"body": "I can build all day but I have never posted on Reddit or X ..."
},
"product": "AutoTraction: an AI marketing co-founder that writes and posts, drafts Reddit and X replies, and pitches press.",
"subreddit_rules": "Self-promotion allowed in comments only if it answers the question and is disclosed."
},
"questions": {
"is_request": {
"type": "noul",
"instructions": "Is the author of `thread` asking, directly or implicitly, for help or a tool to do their marketing?",
"criteria": {
"true": "They describe a marketing problem and would plausibly welcome a suggestion.",
"false": "They are venting, sharing a result, or discussing something else."
}
},
"mention_allowed": {
"type": "noul",
"instructions": "Under `subreddit_rules`, may a reply to `thread` name `product` in one disclosed sentence after answering the question?",
"criteria": { "true": "The rules permit it in this context.", "false": "The rules forbid it or the context makes it promotional." }
},
"angle": {
"type": "choice",
"instructions": "Which reply angle best serves the author of `thread`?",
"criteria": {
"routine": "Give them a small daily routine they can keep up.",
"where_to_post": "Tell them which channels fit a solo builder and why.",
"replies_over_posts": "Explain why replying under big accounts beats posting to no followers.",
"none": "No reply would help."
}
}
}
}
And the response, with the parts we branch on:
{
"answers": {
"is_request": { "type": "noul", "noul": 0.93 },
"mention_allowed": { "type": "noul", "noul": 0.81 },
"angle": {
"type": "choice",
"choice": "routine",
"probabilities": { "routine": 0.62, "where_to_post": 0.27, "replies_over_posts": 0.09, "none": 0.02 },
"confidence": 0.58
}
},
"usage": { "input_tokens": 412, "output_tokens": 0 }
}
The policy sits in code, not in the model. If is_request is above 0.85, the thread goes to the reply queue. If mention_allowed is below 0.7, the brief to the writer says "help only, do not name the product", because a helpful reply with no plug is what earns the account history that lets a later plug survive. The angle's confidence of 0.58 tells us two angles were close, so the writer gets both and picks. The chat model then writes two or three sentences. It never had to read the subreddit rules, and it never had to decide anything.
Example: what does this cold-email reply mean?
Outreach replies are the highest-stakes decision in the system, because a wrong move here lands in a real person's inbox with the founder's name on it. This is a Choice, and the criteria are where the work is. Vague labels get vague probabilities; concrete descriptions get sharp ones.
"reply_intent": {
"type": "choice",
"instructions": "What does `reply.body` mean as a response to `outreach.body`?",
"criteria": {
"interested": "Asks a question, requests a demo or pricing, or says yes to the call.",
"objection": "Engages but pushes back on price, timing, fit, or trust.",
"not_now": "Polite decline that leaves the door open, or asks to be contacted later.",
"wrong_person": "Says they are not the right contact, with or without naming who is.",
"out_of_office": "Automatic reply; no human has read the email.",
"unsubscribe": "Asks to stop, in any tone, including one-word replies like 'no' or 'remove'."
}
}
Then the confidence bands decide who acts. This is the same three-tier rule that runs through the whole product: things the autopilot does alone, things it drafts for approval, and things it hands to the founder.
| Confidence | Who acts | Example |
|---|---|---|
| Above 0.9 | Autopilot | unsubscribe at 0.97: suppress the contact and stop the sequence, immediately |
| 0.5 to 0.9 | Autopilot drafts, founder approves | interested at 0.74: draft the booking reply, ping the founder |
| Below 0.5 | Founder | Split between objection and not_now: show both readings, take no action |
One rule overrides the bands: any answer with meaningful weight on unsubscribe stops the sequence regardless of what won. An unsubscribe is not a preference to be weighed; it is a condition that must not be missed. TypeSafe's own guidance says the same thing: an "any serious violation" rule needs a separate condition, not a weight.
Example: lead scoring you can re-weight without re-running
The old way to score a lead with an LLM was "rate this lead from 1 to 10". The number came back looking precise and meant nothing, and if you changed what you cared about you had to re-run every lead. With Jev we score four dimensions separately, store the raw scores, and combine them in code.
"problem_fit": {
"type": "score",
"instructions": "How closely does `lead.summary` describe the problem `product` solves?",
"criteria": [
"Different problem entirely",
"Adjacent; they might benefit but do not say so",
"Describes the problem in their own words",
"Describes the problem and has tried to solve it already"
]
},
"urgency": {
"type": "score",
"instructions": "How urgent is the marketing problem for `lead`?",
"criteria": [
"No timing signal",
"Mentions it as a someday concern",
"Working on it now",
"Has a launch, deadline, or runway pressure attached"
]
},
"authority": { "type": "score", "instructions": "Can `lead` decide to buy a tool like `product` alone?", "criteria": [ "No", "Influences the decision", "Decides alone" ] },
"budget_signal": { "type": "score", "instructions": "Is there evidence `lead` pays for software like `product`?", "criteria": [ "None", "Uses free tools only", "Pays for at least one comparable tool", "Mentions a budget or a paid stack" ] }
Each score comes back as a number on its levels, with probabilities and confidence. The rank is then one line of arithmetic, and the weights are a setting, not a prompt:
warmth = 0.40 * fit + 0.30 * urgency + 0.15 * authority + 0.15 * budget
When a founder tells us "I only care about people who are launching this month", we move weight onto urgency and re-sort the queue. Nothing is re-inferred, because the evidence and the question meanings have not changed. The same four numbers also make honest features if we ever want to train something classical on which leads actually replied.
Why not just use the LLM you already have?
We did, and it worked, in the sense that the product ran. Three things pushed us to move the decisions out.
Latency inside loops. A scan job looks at thousands of candidate threads and mentions a day and decides on each one. A chat model at a second or two per decision turns "check the alerts" into a job that takes longer than the interval it runs on. A single-pass decision model does not have that problem. We have not published our own timing numbers yet and will not until we have a month of production data, but the shape of the problem is what matters: decisions happen a thousand times more often than writing.
Calibrated numbers you can threshold. When you ask a chat model for a confidence score, you get a number it wrote, and it tends to write 0.8. Jev's probabilities are the model's actual distribution over the options, and its confidence is a measure of how concentrated that distribution is. That is what makes the three-tier rule work: the bands mean something stable, and we can tune them against outcomes instead of against a model's mood.
Separating judgement from generation. The most useful thing about this design is what it did to the writer. The writer model used to receive a long prompt that included the rules, the decision criteria, and the request to write. Now it receives a brief: this thread, this angle, name the product or do not, two sentences. Shorter prompts, fewer refusals, less drift, and a much smaller blast radius when a prompt change goes wrong.
What stays in code, and what stays with the LLM
It would be easy to read this as "Jev decides everything". It does not. The split we landed on:
- Code owns the rules. Posting cooldowns, per-channel rate limits, the requirement that a Reddit plug carries a disclosure, the mutual-follow check that X's ranker rewards, the founder's approval settings. None of these are judgements; they are facts, and a model should never be asked to remember a fact code already knows.
- Jev owns the judgements. Anything where the question is "what does this mean" or "is this the kind of thing that", over state code has already assembled.
- The LLM owns the writing. Replies, posts, pitches, the weekly plan. It also handles the small number of cases Jev flags as low confidence and high stakes, where a slower reasoning pass is worth its cost before a human sees it.
The state we hand to Jev is always observed facts plus the founder's stated preferences. We keep inferred things, like "this lead is warm", in a separate column from observed things, like "this lead replied", so a stale inference can never be mistaken for evidence when the situation changes.
What we would tell another founder building this
Start from the branch points in your own code, not from the model. Write down every if that currently depends on parsing model output. Each one is a question. Ask the ones that share state together in a single request. Put the possible answers in the criteria, in concrete language, and always include a "none of these" option so the model is never forced to pick something that does not fit. Keep thresholds in code and start them conservative. Then run it for a while on real traffic with the autopilot in draft-for-approval mode, and only widen the autopilot band where the approvals are boring.
That last step is the one that matters. A calibrated probability is a promise about the interface, not a guarantee of truth. It has to be checked against what actually happened, on your data, before anything posts without a human.
Frequently asked questions
Is Jev a large language model?
No. Jev understands natural language input but does not generate text. It returns typed answers with probabilities in a single pass, which TypeSafe calls a System One model. You still need a language model to write anything.
Can Jev write my posts or replies?
No. Jev decides; it does not write. In AutoTraction, Jev decides whether and how to reply, and a language model writes the reply from that brief.
How much does Jev cost?
At launch, TypeSafe priced Jev at $0.042 per million input tokens with output tokens free. It is available directly from TypeSafe's API and through Vercel AI Gateway and Cloudflare Workers AI.
What are Noul, Choice and Score?
They are Jev's three question types. Noul returns the probability that a yes/no condition holds. Choice picks one option from a set you define and returns the probability of each. Score places the input on ordered levels you describe and returns probabilities and a confidence number.
Do I need to know any of this to use AutoTraction?
No. This is how the autopilot works under the hood. As a founder you see decisions as approvals, alerts and a ranked queue, and you can tighten or loosen how much runs on autopilot from settings.