Blog Article
Jev vs Laya vs Kev: Comparing AI Decision Models for Enterprise Agents
Jev, Laya, and Kev are changing how AI agents make decisions. Explore their approaches, benchmarks and what decision models mean enterprise AI in Europe.
Jev, which was introduced last week, is now the reference point for every new decision model built in the last 96 hours, but its position at the top of the decision-making benchmark does not defend itself because Laya arrived in a few days to compete against the published benchmark scores with a simple open-source checkpoint, and Kev’s open release just a few days later shows how an open project can build the system behind the model without even shipping the product. The easiest way to highlight this is that years of research have been lost to a weekend of open-source work. The public perception tells us that rivals can have the same isolated numbers, and mechanics can get copied while their system stays unmatched
For enterprises moving from experimentation to production, the bigger question is how these decision models fit into agentic AI systems that can reason, act, use tools and operate inside real business workflows?
What is Jev?
Jev is the decision-making model TypeSafe AI introduced. Their founder, Diogo Almeida, unveiled it on 14 September 2026, and it is the first public release of what the company defines as a “System One model” designed to make fast and structured judgements that software can use directly. Instead of generating text using tokens, it takes the unstructured state + declared questions and gives type-safe structured values with confidence values, probabilities and choices in a single parallel pass. The latency is reported to be 70-500ms end-to-end and priced at $0.042 per million input tokens, with no output charge being billed.
What is Laya, and how does it compete against Jev?
Laya is an open-source decision-making model from ConvAI Innovations, which was released by Nandakishor Mukkunnoth a few days ago. It was released days after Jev as its first competitor. Its moat checkpoint is where the competitor reports a higher hard-label accuracy than Jev but not better outputs on every metric measured. The model card reports 400 test cases with 2000 decisions.

Training exposure is another parameter to be taken on record, as the specialist was tuned on the benchmark’s training split with English and multilingual bases, reaching about 30-35% hard label accuracy in the same published suite, which is below the 46% baseline. A successful specialist is never claimed to be evidence of strong zero-shot performance
Laya competes narrowly against Jev but credibly so for an enterprise that has more than one workflow. The results do not matter enormously, as the customer needs your workflow to run and have a designated intelligence system, but a benchmark specialist must have the absolute right to make decisions on customer data it has not witnessed.
For enterprise teams evaluating different AI stacks, the decision is ultimately tied to the workflow, data, risk profile, and production requirements. Incresco’s AI & Agentic Technologies approach similarly starts with the use case before selecting the technology stack.
How does Kev now behave toward Jev and Laya?
Kev is the decision-making model from Jared Palmer’s open-source decision tree that consists of 3 Qwen3.5-based checkpoints (0.8B, 4B and 9B) with complete training code, frozen suites of evaluation and a server that performs on TypeSafe’s own /v1/systemone API contract. Where Laya competes with Kev is on an open product-building aspect, but Kev competes as proof to show that the decision-making architecture is reproducible with weights, recipe and evaluation that are published under Apache 2.0, and a 0.8B model trains in 20 minutes on an H100 project. Kev ran Jev on its own questions, making the benchmark so close, but Jared warns that the comparison is not controlled because there is no disclosure on Jev’s training data

Kev is not a Jev replacement, but it has built its own model card that calls the checkpoint its research prototype and warns against production that can affect people. For an enterprise to experiment on its own data with complete control of the recipe, which most companies love, then Kev is your starting point, but if you aim for production accuracy and calibration today, Jev is your go-to model for its benchmark that still holds the edge, but we’re waiting to see how many new models will be out in this decision-making world to beat Jev.
What can AI-enabled enterprises do with such out of the world decision making models?
There is no speculation on what benefits AI-enabled enterprises would achieve, as they follow directly from the benchmark alone. The three core properties of the model executes the task are: a work decision that is taken in a shot of 70-500ms, no output cost per decision, minimal input costs and a calibrated sense of confidence where a human must step in and take the final call.
Enterprises have enormously used LLMs to ask questions and get the solutions from them but have never let the models decide, so this breakthrough of models released moves AI from an assistant enterprises used to consult TO an infrastructure that decides the outcome for the enterprise, and this is exactly where production value lives.
Hospitality: Routing Guest Requests
When it comes to Hospitality, every message or query from a guest gets routed from the moment it arrives to decide where humans are involved, from basic housekeeping requests to closing deals with corporates.
But with this decision-making model, an enormous number of certain messages are resolved in an instant with these models, and only uncertain cases are handed to a person, making sure to close requests in a jiffy before the service suffers. If a property receives about 2000 messages a day, it spends about $15 per annum on models, as per Jev’s reported pricing, so at that affordable price point, the question starts moving from whether a hotel can afford to have an AI messaging system built to why the messages are not being routed
This maps directly to Incresco’s Luxury & Hospitality technology solutions, which includes AI-enabled guest support, connected service request handling and operational intelligence systems.
Real Estate: Lead and Enquiry Triage
For real estate, Sales decisions become a boom, with enquiries, viewing notes and tenant requests being decided continuously on certain cases and not a tedious morning inbox sweep.
The boom is in the response time: the faster you decide and respond, the faster you get a decision from the customer, and that’s it you have closed a deal, and that decision layer is what makes a real estate company have an enormous sales track record.
For these workflows, Incresco works across PropTech and Real Estate, including property management platforms, integrated communication and CRM systems, transaction management and predictive analytics.
Education: Assessment and Admissions Workflows
And when it comes to education, you get to implement admission triage and decide thousands of applicants’ admission into your institute and how the assessments are built, with auto-scoring the objective and semi-objective responses with multiple-choice, short numeric, rubric-bucket categorisation and route only the ambiguous 10-20% to a human grader or an LLM for actual written feedback.
You also get to form a hybrid grading pipeline where, once triage earns trust, you formalise the split. Jev runs triage and classification on every submission; an LLM (Claude, for example) touches only the subset that needs written feedback. The final grade is always human-released, with an audit trail on everything, and you get to build an adaptive learning path routing for AI-assisted learning.
These use cases align with Incresco’s Education & EdTech solutions, which cover adaptive learning platforms, assessment frameworks, grading systems and AI-assisted tutoring.
Understand when to flip models.
When Technology and B2B companies come into the picture, this is where your unit economics flip from being trivial to having a strategic decision switch to Laya, as Jev’s model fees run to $38,000 a year on published pricing, with 5 million decisions on average taken in a day. So here is where Laya becomes a cost decision and not an ideology, with its open-weight alternative model being chosen according to the curve companies rely upon hosted simplicity at first and then self-hosted economics at volume.
The deeper gains cut across when calibration makes your automation governable. On Kev’s model suites, Jev has automated 70% of the decisions while holding the errors at or under 5%, so that accuracy level is what an enterprise can take to a board and showcase their defined share of work that has been removed from human queues, along with the speed the models operate at, at a defined budget, with an audit trail on the rest.
For European enterprises specifically, this also connects with the wider question of moving from AI adoption to production deployment. Read Incresco’s analysis of AI adoption across Europe for the broader enterprise context.
That is what AI-enabled enterprises mean to Incresco: not a prototype or a chatbot on the website, but a decision layer inside the operation, making the routine calls instantly so people only touch the ones that matter.
To build that layer in your enterprise, book a discovery call with Incresco. We will evaluate the right decision systems for your workflows and build the agentic solution around them, with measurable ROI.
Let us know what you think of these models - tag us on X and LinkedIn and share what you are building :)