🔍 Read the full analysis: A Decision-Model Playbook: 24 Ways To Use Jev In AI on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A Sept. 29 playbook from Thorsten Meyer outlines 24 proposed uses for Jev, a tool that returns typed answers to narrow questions so software can route routine decisions. Meyer says three publishing uses are already live, 12 are strong fits, seven need measurement and two are poor fits; those classifications and performance figures come from his own account, not an independent evaluation.
Jev takes a text or JSON state and a set of typed questions, then returns structured answers that software can use to make decisions. The source describes three answer types: a yes-or-no probability, a choice among options with probabilities and confidence, and a score on ordered levels. Meyer says Jev does not write, summarize or extract material. A call carries the input and questions together and, according to the playbook, takes about 0.3 to 0.9 seconds.
Meyer reports that three publishing applications are live. One scans article and site pairings for relevance; another checks whether an article is in English; a third classifies topics when a primary model errs. He says an overnight scan of 78,889 articles cost $2.01, found 1,576 non-English items and fixed 1,553. He also reports 89% agreement with a frontier model for the classifier, rising to 97%–99% when Jev’s confidence was at least 0.8. These are the author’s results; the source does not provide an independent audit.
The playbook sorts 24 ideas into three live uses, 12 strong fits, seven cases needing measurement and two poor fits. For each, it gives a question and a rule for acting on the result. Examples in the supplied material include checking whether required disclosures appear on pages and sorting comments into acceptable, spam, abusive or off-topic categories. Some publishing proposals, including detecting thin sourcing and improving headline quality, are marked for measurement first.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Confidence Can Reduce Review Work
The proposed value is high-volume triage: software can act on straightforward cases and send uncertain ones to a person or a more capable model. That approach could make checks practical across large collections of content, while retaining a review path for ambiguous results. The playbook stresses that the application code sets the action; Jev supplies an answer and confidence estimate.
Meyer’s examples also set limits on the case for adoption. A task can be narrow and inexpensive but still lack a demonstrated problem: his duplicate-story test found no duplicates in a canary, so he labels it a poor fit for now. The framework asks teams to measure whether an existing rule fails before replacing it, which makes the proposed value contingent on evidence from each operation.
The Four Tests Behind Jev’s Use Cases
Meyer says Jev should be used only when four conditions are met: high volume, a narrow question, low-cost errors or a route for uncertain cases, and a visibly failing heuristic. If a keyword rule works, his advice is to keep it. For promising cases, he proposes replaying 300 to 500 past decisions, comparing results overall and by confidence band, and examining 20 disagreements to determine which answer was right.
His suggested rollout is to wire the tool in only if the high-confidence band reaches 95% accuracy, keep a separate feature flag off by default, test on 5%–10% of units, then expand. The playbook reports a measurement on a 31-topic classification in which Jev agreed with a frontier model 97%–99% of the time above 0.8 confidence and 42% below 0.5. The material does not specify the sample size or enough evaluation details to establish how broadly those figures generalize.
““Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.””
— Thorsten Meyer, playbook author
How Far the Reported Results Extend
The supplied material is an account by the tool’s author and does not include an independent test, detailed methodology, or a full breakdown of the 24 cases beyond the excerpt provided. General accuracy, performance outside publishing and results on other organizations’ data remain unclear. The source also does not state the evaluation sample size behind the 31-topic agreement figures or provide confidence intervals.
For the seven cases tagged “measure first,” the source says the existing heuristic’s failure has not been proven. It does not establish whether the proposed checks improve outcomes after measurement. The two poor-fit cases are not identified in the supplied excerpt.
Measure Before Expanding Jev
The playbook’s next step for teams considering a use is to replay 300 to 500 real past decisions, inspect disagreements and verify performance in the high-confidence band. Where a use qualifies, Meyer recommends a feature flag and a limited canary before broader rollout. The source does not announce a separate product launch or give a schedule for additional deployments.
Key Questions
What does Jev do?
Jev returns structured answers to typed questions about text or JSON. Its outputs include yes-or-no probabilities, choices with confidence, and scores on ordered levels; software decides what action follows.
How many of the 24 proposed uses are already running?
Meyer says three are live in his publishing operation. He classifies 12 as strong fits, seven as needing measurement and two as poor fits.
What evidence does the playbook report for Jev?
Meyer reports results from his own publishing operation, including a scan of 78,889 articles for $2.01 and agreement figures on a 31-topic classification. The supplied source does not provide an independent evaluation or enough methodological detail to judge how broadly those results apply.
When does Meyer recommend using Jev?
His four conditions are high volume, a narrow question, low-cost errors or a review route for uncertain cases, and evidence that the current heuristic fails. He recommends measuring performance on past decisions before rollout.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
