Probabilistic Decision Primitives Under Distribution Shift: An Empirical Benchmark of JEV
Evaluating JEV, static heuristics, and general LLMs as software decision primitives for dependency-update automation.
“Can a general-purpose probabilistic decision primitive safely extract and retain operational decision signal from software-change metadata under distribution shift?”
We evaluated whether a general-purpose probabilistic decision primitive can extract and retain useful decision signal from software-change metadata under distribution shift. Using dependency updates as the concrete test case, JEV (TypeSafe's System One decision model) was compared against a deterministic static policy and a strong general LLM (DeepSeek Flash) on identical pre-merge state. On a 1,102-case in-distribution benchmark (551 breaking / 551 control), JEV showed substantially stronger ranking quality (AUROC 0.851) than static rules (0.602) and DeepSeek Flash (0.585). On an independently constructed 185-case out-of-distribution (OOD) set (100 breaking / 85 control, 103 independent repositories), JEV retained the highest AUROC among the three (0.605, 95% CI 0.525–0.685), but absolute performance fell sharply. The critical operational finding is negative: JEV's frozen in-distribution threshold did not transfer safely—the 0.62 threshold yielded 30 auto-merges at 50.0% precision with 15 unsafe merges. These empirical results demonstrate that while structured decision models can capture strong ranking signal, probability calibration and threshold stability remain unresolved under domain shift.
JEV ranking quality on 1,102 pre-merge dependency update cases (vs 0.602 static, 0.585 LLM)
Retained highest ranking on 185 independent out-of-distribution cases across 103 repos (95% CI: 0.525–0.685)
Frozen in-distribution threshold produced 15 unsafe auto-merges out of 30 total actions
Total evaluated cases under strict pre-merge isolation with zero post-merge telemetry leakage
| System Evaluated | Model Class | In-Dist AUROC (n=1,102) | OOD AUROC (n=185) | OOD Prec @ 0.62 | Unsafe Merges |
|---|---|---|---|---|---|
| JEV (System One) | Probabilistic Choice | 0.851 | 0.605 | 50.0% | 15 / 30 |
| Static Rule Policy | Deterministic Heuristic | 0.602 | 0.551 | 41.2% | 20 / 34 |
| DeepSeek Flash | General LLM Prompt | 0.585 | 0.514 | 44.4% | 15 / 27 |
01 / Conceptual Framing: Separating Three Distinct Capabilities
In AI decision systems, three fundamental questions are routinely conflated: (1) Decision ranking quality: Does the model score safe updates above unsafe ones? (2) Probability calibration: Do the returned numeric probabilities retain their statistical meaning on unseen distributions? (3) Safe operational automation: Does a fixed confidence threshold enforce safe behavior in production without human intervention? An automated system can exhibit competitive discriminative ranking while catastrophically failing calibration and threshold safety.
02 / Experimental Setup & Hermetic Pre-Merge Isolation
To guarantee empirical rigor, all three systems (JEV, Static Rules, and DeepSeek Flash) evaluated identical pre-merge state: ecosystem, package manager, semantic version diff, PR title/body, manifest changes, release notes, and pre-merge CI status. Strictly zero post-merge telemetry (merge status, subsequent reverts, error logs) was accessible. Furthermore, models were not provided repository source code or API dependency graphs, isolating whether change metadata alone carries transferable decision signal.
{
"task": "DEPENDENCY_UPDATE_DECISION",
"allowed_actions": ["AUTO_MERGE", "HOLD", "REQUIRE_REVIEW"],
"pre_merge_state": {
"repository": "isolated-repo-id",
"ecosystem": "npm",
"dependency": "package-name",
"version_transition": { "from": "1.4.2", "to": "1.5.0", "semver_type": "minor" },
"manifest_diff": "+ package-name@1.5.0\n- package-name@1.4.2",
"ci_status_pre_merge": "SUCCESS",
"has_release_notes": true
},
"leakage_boundary": {
"post_merge_facts_included": false,
"source_code_index_included": false
}
}03 / Empirical Findings: Discriminative Signal vs. Distributional Decay
On in-distribution evaluations (1,102 cases), JEV demonstrated strong discriminative capacity, achieving an AUROC of 0.851—decisively outperforming deterministic rules (0.602) and general-purpose LLM prompting (DeepSeek Flash at 0.585). However, when exposed to 185 out-of-distribution cases across 103 unseen repositories, JEV's AUROC dropped to 0.605 (95% CI: 0.525–0.685). While it remained the top-ranked system, the discriminative advantage over simple heuristics compressed drastically.
04 / Operational Failure: The Frozen Threshold Hazard
The critical finding of the benchmark is operational. In production automation, teams fix confidence thresholds (e.g. threshold = 0.62) to determine whether a pull request can bypass human review. Under distribution shift, this frozen threshold collapsed: out of 30 auto-merge executions, exactly 15 were breaking updates—a 50.0% precision rate equivalent to a coin toss. This empirical outcome establishes that static confidence thresholds cannot be safely frozen across shifting software codebases.
# Clone and reproduce the benchmark results locally
git clone https://github.com/scarif-labs/jev-software-decision-benchmark.git
cd jev-software-decision-benchmark
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python run-eval.py --dataset data/ood-185.json --model jev-latest05 / Implications for Autonomous Software Engineering
For software engineers building agentic workflows and automated dependency systems: change metadata carries genuine predictive signal, but autonomous action requires dynamic calibration rather than static probability thresholds. Autonomous code operations must implement adaptive conformal prediction, rollback telemetry, and graduated human oversight when operating outside their training distribution.
- 01.Ranking signal does not equal operational safety: high AUROC (0.851) does not prevent unsafe actions if calibration decays under distribution shift.
- 02.Frozen confidence thresholds are hazardous: an in-distribution threshold of 0.62 resulted in a 50.0% failure rate (15 unsafe merges) on out-of-distribution repos.
- 03.Domain-specific decision primitives outperform general LLMs in ranking: JEV outperformed DeepSeek Flash across both in-distribution (0.851 vs 0.585) and OOD (0.605 vs 0.514).
- 04.Software automation requires adaptive uncertainty estimation (such as conformal prediction) before autonomous write actions can be safely delegated.
@misc{scariflabs2026jevbenchmark,
title={Independent Benchmark of JEV for Dependency-Update Automation Under Distribution Shift},
author={Lawrance, Alen and PS, Harikrishnan and Prakash, Vishnu},
year={2026},
publisher={Scarif Labs Research},
url={https://github.com/scarif-labs/jev-software-decision-benchmark},
howpublished={\url{https://scariflabs.com/lab/jev-software-decision-benchmark}}
}Applied in our production systems & services
Real-time collaborative infrastructure so distributed teams work together without lags, data loss, or server bottlenecks.
A custom operational sync engine designed for high-concurrency environments. Solves multi-user state synchronization and offline convergence with zero data loss.
Production System · HorusTurns any Android phone into a real Alpine Linux environment to run coding agents locally or SSH in over USB/Wi-Fi.
A complete Alpine Linux userland engineered for Android with no root required. Run Claude Code, Codex, and zsh directly on the phone, or pair over SSH to use your phone as a private edge compute node instead of paying for cloud VMs.
Studio Service · Custom Engineering & SystemsSoftware Development
Custom software engineered with architectural rigor and resilient foundations.
Studio Service · AI-Native Products & AgentsAI Product Development
Design and engineering for AI-native applications, agents, and intelligent workflows.
Interested in empirical AI evaluations or autonomous agent architecture?
We design, evaluate, and productionize intelligent systems and reliable software architectures.