Blog

/

Guides

/

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13: Can You Run It Locally? Laya Setup Guide

Jev 1.13 makes structured decisions for your apps. See Jev and local Laya play Tetris, then set up Laya on your own computer with Python.

Jev 1.13: Can You Run It Locally? Laya Setup Guide
Alex Shapiro
Alex Shapiro
Calendar icon

September 25, 2026

Table of Contents

Jev 1.13 evaluates text and answers structured questions your application defines. Give it a support message and a list of departments, and it can choose where to route the request. If you want to run Jev locally, the official sources reviewed for this guide provide no downloadable weights. For local decision inference, this guide uses Convai Innovations' Laya, a separate model with a Jev-compatible interface. TypeSafe model documentation · Laya repository

In this guide, you'll learn how Jev 1.13 works and what changes when you use Laya locally. Our Tetris demo compares the two deployment paths. The macOS and Linux walkthrough then uses laya==0.3.20 in an external Python process to route a payment complaint. Jump to the Laya installation steps to start the setup.

Verification note: We checked Laya's published package, loader source and checkpoint files on September 25, 2026. We installed laya==0.3.20 without its inference dependencies and confirmed that it imports. Model loading, predictions and disconnected execution remain untested.

What is Jev 1.13?

Jev is TypeSafe's first System One model, a model that evaluates application state and returns typed decisions. Because the application defines the questions and allowed answers, Jev can choose a route or score a condition without writing a response that your code must parse. TypeSafe's System One documentation

The state is the material being evaluated. For a support workflow, it could be a JSON object containing this message: My card was billed twice for the same order. Please reverse the extra charge. Keep the customer's text in the state and the routing rules in the question. State documentation

The following examples describe the interface, not measured model outputs.

PrimitiveQuestion about the messageWhat the application receives
ChoiceWhich queue should receive it: payments, product support or account access?A selected option and probabilities across the options
ScoreHow urgent is it on a defined, ordered rubric?A numeric expected score across the rubric, which can be fractional
NoulDoes the customer request a refund?The model's estimated probability that the proposition is true

On September 25, 2026, TypeSafe lists jev-1.13.0 as the current version; jev-latest and jev-preview both resolve to it. Jev 1.13 accepts text, including JSON state. Its documented limits are 64k tokens across a request and 32k for the state plus the longest question. These are input limits; long, irrelevant state can still reduce accuracy. Model specifications · Jev 1.13 limitations

TypeSafe defines calibration across groups of predictions, so even a calibrated model can route an individual payment request to the wrong queue with high confidence. Calibration scope

Use ordinary code to check transaction amounts and permissions. Jev's own limitations page calls out unreliable arithmetic, counting and date calculations, as well as sensitivity to how instructions and evidence are phrased. Jev 1.13 limitations

Can you run Jev 1.13 locally?

There is no official local installation route in the sources reviewed for this guide. TypeSafe serves Jev through POST /v1/systemone, with Python and JavaScript SDKs available. Running its API client on a laptop still sends inference to the service. TypeSafe model list

For local decision inference, download Laya's weights and load them in Python. The installation below runs Laya; it does not download or execute Jev 1.13.

Jev vs Laya: what changes when you go local?

Laya is Convai Innovations' independent model family with downloadable weights and an Apache 2.0 license. Its Jev-compatible interface can make application code easier to adapt, but it does not establish equal accuracy or identical architecture. Laya repository

This comparison uses TypeSafe's model documentation and Laya's published model card.

DetailJevLaya
DeveloperTypeSafe AIConvai Innovations
Deployment describedHosted serviceExternal local Python runtime
Version to identifyjev-1.13.0 for the documented releaseRuntime version plus checkpoint and revision
InputText state, including JSONText state, including JSON
OutputTyped decisions and probabilitiesTyped decisions and probabilities
Local weights in reviewed official sourcesNo official download foundAvailable on Hugging Face under Apache 2.0

If your requirement is to keep decision inference on your computer, start with Laya. If you want a managed service, evaluate Jev on the same labeled examples.

Our Jev vs Laya Tetris demo

We compared cloud-based Jev with local Laya in a Tetris demo posted by @atomic_chat_hq. The post identifies a MacBook Air with 16 GB of memory for the local run. In the footage, Jev's board reaches game over while Laya continues playing. The board can change while a model is choosing its next move, which makes response time relevant to a game loop.

Atomic Chat's recorded Tetris demo. The visible timers are application display values; their timing boundaries are not documented in the post. Open the original post.

ItemConfirmed detailLimit
TaskTwo side-by-side Tetris boards labeled Jev cloud and Laya localGame code, action schema and board synchronization are not supplied
Local hardwareThe post specifies a 16 GB MacBook AirProcessor generation and runtime are not specified
Visible timersOne inspected frame shows Jev at 316 ms and Laya at 48 msA frame is not an average or a latency distribution
Visible outcomeJev reaches game over while Laya continuesThe number of attempts and selection of this recording are unknown
Speed claimThe post says Laya made decisions 11 times fasterThe inspected frame does not reproduce that ratio; underlying logs are unavailable
Model versionsThe footage labels the models Jev and LayaExact Jev version, Laya checkpoint and revisions are not identified

Without the measurement code, we cannot separate model inference from network time, assign the cloud delay to network latency alone, or use the game outcome to establish an accuracy winner across tasks.

For your own comparison, record the same states and questions for both models. Measure local loading separately from warm predictions, and distinguish full API response time from any server-reported inference time. Keep every attempt, then compare decision accuracy alongside latency.

Laya checkpoints and computer requirements

The Laya model card describes these three options; its language coverage and task results are developer reports.

CheckpointModel sizeDefault input budgetIntended starting point
convaiinnovations/laya421M parameters512 tokensEnglish decisions
convaiinnovations/laya-multilingual322M parameters1,024 tokensMultilingual inputs; the developer reports coverage of more than 100 languages
convaiinnovations/laya-typed-decisions421M parameters1,024 tokensInvoice processing, security incidents, customer service and agent-trace observability

The typed-decisions checkpoint was fine-tuned on the training split of the developer's benchmark for these workflows. In that evaluation, the base checkpoints scored below a baseline that always chose each question's most common answer. Developer's evaluation

For the walkthrough below, use the English checkpoint with Python 3.10 or newer. The package declares dependencies on PyTorch, Transformers, Safetensors, Hugging Face Hub and NumPy. CPU is a supported path; NVIDIA CUDA and Apple's MPS backend are additional device options. The steps below use CPU to keep device configuration explicit. Package configuration

The English model uses a bidirectional ModernBERT-large encoder and a decision head that scores the supplied options. Multilingual uses mmBERT-base; its documented input limit can be raised to 8,192 tokens explicitly. These budgets include the serialized material the model evaluates, not an unlimited document plus a separate question. Long text and large option lists need testing for truncation. Runtime documentation

The English checkpoint's weight file is approximately 843 MB, plus tokenizer and configuration files. That is a download size; estimating RAM also requires accounting for PyTorch, intermediate tensors and other running applications, none of which was measured in this review. Our Tetris post's 16 GB machine is an observed setup, not a measured minimum. Checkpoint file list

For Apple Silicon, the independent laya-mlx port offers a separate MLX runtime and converted checkpoints. Its package requires Python 3.11 or newer and macOS 14 or newer; because it uses a different execution backend, timing results from the PyTorch guide below would not describe the MLX setup.

How to run Laya locally with Python

This route runs Laya in a Python process on your computer. We have not verified native Laya catalog support or an end-to-end Laya integration inside Atomic Chat.

1. Create a Python environment

These shell commands target macOS or Linux with Python 3.10 or newer installed. On Debian or Ubuntu, an ensurepip error during environment creation usually means the matching Python venv package needs installing.

python3 --version
mkdir laya-local
cd laya-local
python3 -m venv .venv
source .venv/bin/activate
python -m pip install "laya==0.3.20"
python -I -c "import laya; print(laya.__version__)"

The last command should print 0.3.20. It checks the package import, not a model prediction. Installation also downloads dependencies; their disk usage is separate from the model weights.

2. Download one pinned checkpoint

Save this as download_model.py. It downloads the English checkpoint at the revision inspected for this guide. The filter includes its weights, tokenizer and configuration files, leaving out the multilingual and typed-decision checkpoints.

from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="convaiinnovations/laya",
    revision="55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851",
    local_dir="models/laya",
    allow_patterns=[
        "rl_agent_config.json",
        "model.safetensors",
        "tokenizer/*",
        "encoder/*",
    ],
)

Run it while connected to the internet:

python download_model.py

Keep models/laya intact. The published loader needs its weight file, rl_agent_config.json, encoder/config.json and tokenizer files together.

3. Run the first decision

Save this as first_decision.py in the same folder. The question asks for a queue; the program prints the selected label and the probability assigned to each candidate.

import json
import laya

agent = laya.load("./models/laya", device="cpu")

state = {
    "message": (
        "My card was billed twice for the same order. "
        "Please reverse the extra charge."
    )
}
questions = {
    "queue": {
        "type": "choice",
        "instructions": "Which support queue should handle this message?",
        "criteria": {
            "payments": "Billing errors, charges and refunds.",
            "product_support": "Product bugs or failures.",
            "account_access": "Login and password problems.",
            "manual_review": "Unclear requests or none of the other queues.",
        },
    }
}

result = agent.predict(state, questions)
answer = result["answers"]["queue"]
print("Selected queue:", answer["choice"])
print(json.dumps(answer["probabilities"], indent=2))
python first_decision.py

For this message, payments is the intended classification. A successful model call produces a selected key and a probability dictionary; it does not guarantee the intended answer. We have not measured this example's inference output.

4. Check a subsequent offline run

After the first prediction succeeds, rerun the script with Hugging Face and Transformers in offline mode:

HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1 python first_decision.py

This disables Hub lookups for the libraries used here. The script loads a local directory and contains no remote inference call. To confirm the entire setup works disconnected, disconnect networking and repeat the command. Offline execution has not been tested for this guide; a missing tokenizer or encoder configuration can still prevent loading.

Keep the virtual environment. Once your prediction succeeds, record its resolved dependencies with python -m pip freeze > requirements-lock.txt so you retain the package versions that worked with your model files; pinning Laya alone does not pin every dependency.

Use Laya to route support requests

Map the returned queue to an internal inbox, and retain the original message for the person handling it. Add labeled examples of account-access problems and product failures to check that the model distinguishes the queues.

Include ambiguous requests and messages that do not belong in any named queue. The model can choose manual_review for these requests, but it may still assign them to the wrong queue with high confidence. Test that behavior on held-out examples from your own workload before automating the routing.

Keep refund authorization in application code. A request classified as payments says where it belongs; transaction records and your refund policy determine whether any money should move. If you want a written customer reply, use a text-generating model after routing. Our local models guide for a 16 GB MacBook covers that separate workload.

Laya also documents an optional MCP server, and Atomic supports MCP tools. We have not verified this end-to-end pairing.

Common Laya installation and output problems

Source: Laya runtime documentation and limitations.

ProblemWhat to check or change
No module named layaActivate .venv and use python -m pip from that environment. Avoid naming your script laya.py.
Download or offline-cache errorRun the download step online, then confirm the local directory contains weights, configuration and tokenizer files. Run the scripts from laya-local.
CUDA or MPS is unavailableKeep device="cpu" for the first run. Configure the matching PyTorch backend before switching devices.
Memory use is higher than expectedLoad a single checkpoint. Router(preload=True) can load multiple checkpoints; it is unnecessary for this example.
Long input loses relevant evidenceShorten the state and option descriptions. Check the selected checkpoint's input and decision-head budgets before increasing them.
English Noul answers look wrongThe model card reports sensitivity to false/true option wording. Evaluate an explicit two-option Choice with neutral labels and clear descriptions.
action.act_probability is almost always highThe model card reports that this head has little useful variation. Do not treat it as permission to act.

For Choice and Score, Laya's confidence field is entropy-based; answer_confidence is the largest option probability. For Score, that probability belongs to a rubric level, not to the fractional expected score. Answer-decoding implementation

Laya's developers report that the English and multilingual base checkpoints are overconfident as shipped. Evaluate their probabilities on labeled examples from your own workload before using them to control automated actions. Model-card limitations

Frequently asked questions

Is Laya an open-source version of Jev?

No. Laya is an independent model family from Convai Innovations with Apache 2.0 code and weights. Its decision interface overlaps with Jev's, but it is not an official Jev port or a download of Jev's weights.

Can I run Laya on a Mac?

Yes. The Python runtime supports CPU and MPS device paths, and the independent MLX port targets Apple Silicon. Our Tetris post reports a 16 GB MacBook Air, without specifying its chip or runtime. Use that as evidence of the demo setup, not a minimum-memory specification.

Does Laya need an API key?

No API key is needed for this local Python example. It uses a public, ungated checkpoint. Installing packages and downloading files initially requires internet access. A separately hosted service can have its own authentication requirements.

Can Laya write replies or replace a chat model?

No. Laya returns decisions within the question types you define. A workflow that needs an explanation or a drafted reply still needs text generation. Keep the two calls separate so you can evaluate routing and response quality independently.

Is Laya always faster or more accurate than Jev?

No universal result follows from our Tetris demo. Runtime, input length, hardware and the API route affect response time; the task and checkpoint affect accuracy. Measure both on the workload you plan to deploy.

Which Laya checkpoint should I start with?

Use the English checkpoint for the short English example in this guide. Evaluate the multilingual checkpoint for other languages. If your task matches the specialist workflows in the checkpoint table, evaluate the typed-decisions checkpoint on your own labeled data.

Should you use Jev 1.13 or local Laya?

Use Jev 1.13 if you want TypeSafe's hosted decision model and can send the required state to its API. Choose Laya if running decision inference on your own hardware is the requirement. After your first local prediction succeeds, replace the sample message with labeled cases from your application. Compare the selected queues with your labels before connecting the results to a live inbox. If you also evaluate Jev, send it the same cases so you can compare routing errors as well as response time.

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Best Local AI Video Generators Compared on Quality, Speed and VRAM

Compare local AI video generators on quality, speed and VRAM, with RTX 4090 benchmarks, example clips and an Atomic Chat setup guide.

10/9/26

16 min

Best Local LLMs for 12GB VRAM in 2026

Best Local LLMs for 12GB VRAM in 2026

Eight local LLMs for 12GB VRAM: exact GGUF files, RTX 3080 Ti speed and memory figures, plus Snake and physics tests on an RTX 4070.

9/30/26

15 min

Claude Sonnet 5.5 Alternatives Compared

Claude Sonnet 5.5 Alternatives Compared

Compare Claude Sonnet 5.5 with Opus, Fable and GPT-6 Astra, then choose a local Qwen, Ornith or Bonsai model for your hardware.

9/29/26

13 min

Best Local AI Image Generators in 2026: Apps and Models

Best Local AI Image Generators in 2026: Apps and Models

Compare local AI image generators, with seven models benchmarked on an RTX 5090. See image examples, generation speed, memory use, and license limits.

9/25/26

14 min