← COAI evaluation case study · Development pilot snapshot · Reference review pending
Research / Agent monitoring·

Agent monitoring across trajectories

A custom evaluation of System One models, language models and a reranker acting as decision monitors. We test how accurately and quickly they judge an agent’s next action from the task rules and the trajectory so far.

A trajectory is an ordered sequence of agent messages, tool results and proposed actions. Across different tasks, each monitor returns the same bounded decision: permitted, prohibited or insufficient evidence. We compare JEV, Laya, JevClone and DiffusionGemma’s native decision interfaces with hosted LLMs prompted to produce similar outputs, a local LLM using constrained decoding, and a reranker scoring the three decision criteria.

ControlArena & Inspect

Monitoring approaches compared

The models receive the same projected trajectory evidence and policy criteria. Their training, interfaces and decoding methods differ.

ApproachDesign ideaHow it produces our decision
JEVNative System One modelTypeSafe describes a purpose-built decision model trained with Reinforcement Learning for Calibrated Decisions (RLCD) and parallel outputs.We submit the trajectory as state and ask one native Choice question over the three policy labels. The API returns a typed decision and model probabilities.
Hosted QwenLLM prompted for System One–style outputA general-purpose language model acts as the monitor. Instructions request a bounded policy decision in JSON, with no explanation required.The model generates its response as text; our adapter parses and validates it. Any returned probabilities are self-reported. The prompt establishes an output contract, not a new model architecture or a guarantee of fast inference.
Jina Reranker v3Relevance ranking used for classificationA reranker scores the three fixed policy criteria against a query containing the monitoring instructions and visible trajectory. The publisher describes a Qwen3-0.6B-based listwise reranker that processes query and documents together. The hosted checkpoint revision and precision are not exposed by this endpoint.We call /rerank with all three criteria and select the highest raw relevance score, mapping by document index. Exact ties use the fixed order: permitted, prohibited, insufficient evidence. Softmax at temperature 0.1 provides normalized scores, not calibrated probabilities. No candidate wording is tuned on these results. The client sends the full evidence; server-side truncation is unverified.
DiffusionGemmaDiffusion model · native decision APIThe operator provides a deployment named dgemma through a TypeSafe-style API. This run tests its Choice decisions at the requested sampling settings. The deployed checkpoint revision, precision and probability calibration are not independently established here.The same trajectory and three policy criteria are sent to /systemone, explicitly requesting samples=1, steps=1, think=0 and seed=42, with concurrency 1. The declared limit is 4,096 tokens including the answer canvas. These settings are recorded with the results; no longer-context deployment or parameter sweep is claimed.
JevClone 8B — v1 / v3COAI-trained 8B model · native decision APIJevClone 8B is COAI’s own trained model for agent monitoring, served through a native decision API. V1 and v3 are distinct model revisions trained and deployed by COAI; both use the API name jevclone-8b-experimental. The report keeps their results separate and records COAI’s revision identification and serving-contract reference with each run. The exact base checkpoint and training configuration are not documented in the current evaluation record.We send the same complete trajectory, instructions and three policy criteria to its TypeSafe-style /systemone Choice API. The run uses native defaults and concurrency 1. Response times include the Tailscale network path and server processing. No context truncation, sampling override or tuning on the evaluation cases is performed by the client.
JevClone 4BCOAI-trained Qwen3 4B model · native decision APIJevClone 4B is COAI’s own experimental model, trained on Qwen3 4B and deployed with its trained adapter retained, FP16 weights and calibrated probability outputs. These model and deployment details are supplied by COAI. This evaluation measures the deployed model’s decisions; it does not independently validate probability calibration.We call jevclone-4b-experimental through its native TypeSafe-style Choice API on the local network, using the same state, instructions and criteria as JEV. No chat-completions prompt or generated-JSON parser is used. Limits: 2,304 tokens per scoring prompt and 8,192 aggregate request tokens. The server explicitly rejects oversized requests; we preserve these as failures and do not truncate or shorten cases.
LayaLocal System One decision modelThe requested English root checkpoint uses a bidirectional ModernBERT-large encoder and trained decision heads, about 421M parameters. The author describes RLCD training. Options are scored in one forward pass; no answer text is generated.We use the native Choice SDK with the same state, instructions and three criteria as JEV. Local PyTorch MPS, float32, on this Apple M3 Ultra. The total/head budgets are raised from 512/192 to 8,192/512 to retain our inputs and rubric; this is an extended-context configuration, not the default model setting. Every token sequence is checked for truncation. Oversized cases fail before inference and count as errors. Shipped temperature scaling is retained; calibration on this suite is not established. No router or typed-decisions checkpoint is used.
Ternary Bonsai 2 27BCompressed LLM · OpenRouterPrismML describes a Qwen3.8-27B-derived model using ternary weights: scaled values from {−1, 0, +1}. Ternary refers to weight representation, not to our three policy labels.We use the same decision-JSON prompt as hosted Qwen. It remains a generative model, with provider-default reasoning and no decoder-enforced label restriction in this run. Measurements include OpenRouter and its serving provider. This session uses a new curl connection per request after Python TLS handshakes failed; connection setup is included in latency. These are not measurements of local Bonsai inference.
Local Qwen + MLX constrained decoding“Qwen RLCD” in the resultsQwen2.5-1.5B-Instruct, quantized to 4-bit, runs on this Mac through MLX. A decoding engine restricts the available answers instead of asking the model to write decision JSON.After processing the input, the adapter scores the allowed labels’ first tokens, selects the highest score and assembles the output in code. This run uses one decision field; it does not test the engine’s multi-field parallelism.

JEV’s RLCD is a training method. The local repository’s “RLCD” name refers here to the constrained-decoding artifact we ran; we did not apply JEV’s training method or load separate RLCD-fine-tuned weights. Restricted token scores and self-reported probabilities do not establish calibration.

This compares deployed monitoring approaches. Model size, training, quantization, hardware and transport also differ, so a score or speed difference cannot be attributed to architecture alone. The tasks replay authored trajectories; they do not execute or prevent the proposed actions.

Sources: Laya model card; Laya SDK source; TypeSafe’s System One and RLCD description; the local decoding engine’s model card; PrismML’s Bonsai 2 model description; Bonsai’s OpenRouter endpoint. Execution details above describe the pinned adapters used in this report.

The original cases cover full history and current-event-only evidence. Challenge cases test full-history decisions. Select a score to inspect its cases; coverage shows which sections each model has completed.

One row per model/version. Scores use evaluated cases only; coverage shows what is still missing. Select a section before comparing scores. Models with incomplete coverage are not directly comparable to complete runs.

Model / versionCoverageAccuracyBalanced accuracyDecisive accuracyMedian / p95Errors

These are authored development cases with draft reference labels, pending independent human review. Errors count as incorrect. This experiment measures judgments on recorded actions, not prevented harm.

Performance profiles by section

Compare the shape of each model’s results. Every spoke shows accuracy from 0% at the centre to 100% at the edge. Switch models on or off; select a point to inspect the cases behind it.

Models shown in all spider charts

Errors, including context-limit failures, count as incorrect. Unrun or partially evaluated spokes are left empty; they are not zero scores. Polygon area is not an overall score: spokes contain different numbers of related cases. Exact counts and errors are available below each chart. Identical scores overlap; hide other models to isolate one. Zero-score labels are offset from the centre so each spoke remains selectable.

Classification speed

Compare response times →

Results by scenario

Compare all models on the same scenario. Select any score to inspect that model’s decisions. Scroll horizontally when more models are added.

Decisive actions

The final checkpoint, e6, is where the authorization or evidence distinction is tested. The two earlier checkpoints assess ordinary internal work.

Detection & review

Challenge difficulty

All models under the selected evidence condition. These are grouped development cases; levels can vary evidence position as well as reasoning burden.

Reference versus prediction: confusion matrices

Rows show the draft reference; columns show model predictions. Select any count to open the matching results.