🔍 Read the full analysis: 24 Jev Use Cases For Building And Applying Decision Models on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer published a map of 24 proposed Jev uses across publishing, commerce, software, business operations and the home. He says three are running in his publishing operation, 12 meet his fit test, seven need more measurement and two are poor fits; the reported results are his own, and the source does not provide the remaining use case details.
Thorsten Meyer has published a map of 24 uses for Jev, a tool he describes as returning typed answers to questions about text or JSON so software can make decisions. In his September 29 account, Meyer says three checks are live in his publishing operation, which has processed about 90,000 decisions, while 12 other uses meet his stated fit test. The figures and performance results are Meyer’s own reports; the supplied source does not include independent validation.
Meyer groups the proposed applications across publishing, commerce, software, business operations and the home. He classifies 12 as strong fits, seven as needing measurement and two as poor fits. His article says each use case is framed around a question, a question type and a rule for acting on the answer. The source material provided here includes detail on the production examples and the first six publishing cases, but cuts off during the commerce section; it does not give the full list of 24.
In Meyer’s description, Jev accepts a state and typed questions, then returns an answer that code can use directly. The answer types include a yes probability, a choice among options with probabilities and confidence, or a score on ordered levels. He says a call typically takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens. Those operating figures are presented by Meyer, not as independently audited measurements.
The three live publishing uses are a relevance check, a language check and a classifier fallback. Meyer reports that a one-night scan covered 78,889 articles for $2.01; the language check found 1,576 non-English articles, of which 1,553 were fixed. For a 31-topic classification, he reports 89% overall agreement with a frontier model and 97% to 99% agreement when Jev’s confidence was at least 0.8. The source does not state the size or method of that comparison in the supplied passage.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The proposal matters to teams that handle large volumes of routine decisions. Meyer argues that a low-cost check can cover cases that might otherwise be missed by a manual or keyword-based process, while ambiguous cases continue through an existing workflow or go to a person. In publishing, that could mean flagging language, relevance or topic classification across many articles.
The practical case depends on more than speed or cost. A tool can add little value if a simple existing rule already works, or if errors carry serious consequences. Meyer’s own example of a same-event duplicate detector was tagged a poor fit after a canary test found zero duplicates. That result illustrates his central operational point: measure a real failure before adding another automated decision step.
Meyer’s Four-Part Fit Test
Meyer says a Jev use should meet four conditions: high volume, a narrow question without multi-step reasoning, low-cost errors or a path for uncertain cases to receive more capable review, and a visibly failing heuristic. He advises teams to establish the last condition with data rather than assume a problem exists.
His proposed validation process is to replay 300 to 500 past decisions, compare results overall and by confidence band, and review 20 disagreements. Meyer says to integrate a use only where the high-confidence band reaches 95%, then put it behind a separate feature flag, start with 5% to 10% of units and expand gradually. These are his recommendations, not a reported industry standard.
Among the publishing cases shown, Meyer labels disclosure detection and comment moderation strong fits. He says disclosure checks could catch affiliate or free-product statements that a regular expression misses, with misses sent for human review. For moderation, he proposes automatically approving clearly acceptable comments or hiding clearly spam or abusive comments at confidence of 0.9 or higher, with uncertain cases queued. Thin-source detection, product matching in roundups and headline quality are marked for measurement first.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer
Independent Results Remain Unverified
The supplied material is an account by Meyer and does not include an independent evaluation, full methodology or dataset for the reported results. It is also incomplete after the start of the commerce section, so the questions, rules and evidence for many of the 24 use cases cannot be assessed from this source alone.
Several performance details remain unspecified in the excerpt, including how the 31-topic comparison was conducted and how the 20 disagreements are selected for review in the proposed testing process. The source also does not describe the deployment environment, error rates for each live check, or the impact of the 23 language-check cases that were found but not reported as fixed.
Measure Before Expanding Deployment
Meyer’s recommended next step for teams considering a use is to test it against 300 to 500 past decisions, examine disagreements and confirm that high-confidence answers meet his 95% threshold. If a use passes, he proposes a feature flag and a small canary rollout before broader use, with uncertain cases routed for review.
The supplied article excerpt does not identify a release schedule or provide further milestones for Jev. More detail on the other proposed applications, and independent evidence about the live checks, would help readers judge how broadly Meyer’s results apply.
Key Questions
What did Meyer announce?
He published a map of 24 proposed Jev uses across five areas, and reported three publishing checks running in his operation.
What does Jev return?
According to Meyer, it returns typed answers, such as a yes probability, a choice among options or a score. Software can use the result to branch without interpreting a prose response.
Are the reported results independently verified?
The source provides Meyer’s account and does not include independent verification or the full methods behind the reported performance figures.
How does Meyer suggest testing a use case?
He recommends replaying 300 to 500 real past decisions, checking performance across confidence bands and reviewing disagreements before a gradual rollout.
Does the supplied source describe all 24 uses?
No. The supplied text gives detail on three live publishing uses and six publishing proposals, then cuts off during the commerce section.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
