Probabilistic Decision Primitives Under Distribution Shift: An Empirical Benchmark of JEV

Evaluating JEV, static heuristics, and general LLMs as software decision primitives for dependency-update automation.

Empirical BenchmarkIdentifier: SCARIF-RES-EXP-05
GitHub Repository ↗
Research Investigators
✦Alen Lawrance(Scarif Labs)✦Harikrishnan PS(Scarif Labs)✦Vishnu Prakash(Scarif Labs)
Published: 2026-09-25
Revision: 2026-09-28
// core_questionTarget Workflow: Autonomous Dependency Merging

“Can a general-purpose probabilistic decision primitive safely extract and retain operational decision signal from software-change metadata under distribution shift?”

Abstract

We evaluated whether a general-purpose probabilistic decision primitive can extract and retain useful decision signal from software-change metadata under distribution shift. Using dependency updates as the concrete test case, JEV (TypeSafe's System One decision model) was compared against a deterministic static policy and a strong general LLM (DeepSeek Flash) on identical pre-merge state. On a 1,102-case in-distribution benchmark (551 breaking / 551 control), JEV showed substantially stronger ranking quality (AUROC 0.851) than static rules (0.602) and DeepSeek Flash (0.585). On an independently constructed 185-case out-of-distribution (OOD) set (100 breaking / 85 control, 103 independent repositories), JEV retained the highest AUROC among the three (0.605, 95% CI 0.525–0.685), but absolute performance fell sharply. The critical operational finding is negative: JEV's frozen in-distribution threshold did not transfer safely—the 0.62 threshold yielded 30 auto-merges at 50.0% precision with 15 unsafe merges. These empirical results demonstrate that while structured decision models can capture strong ranking signal, probability calibration and threshold stability remain unresolved under domain shift.

// benchmark_telemetry
In-Distribution AUROC
0.851

JEV ranking quality on 1,102 pre-merge dependency update cases (vs 0.602 static, 0.585 LLM)

OOD AUROC
0.605

Retained highest ranking on 185 independent out-of-distribution cases across 103 repos (95% CI: 0.525–0.685)

OOD Precision @ 0.62
50.0%

Frozen in-distribution threshold produced 15 unsafe auto-merges out of 30 total actions

Pre-Merge State Pairs
1,287

Total evaluated cases under strict pre-merge isolation with zero post-merge telemetry leakage

// comparative_evaluation
Comparative Evaluation Across In-Distribution and Out-of-Distribution CohortsHermetic pre-merge isolation (Zero post-merge leakage)
System EvaluatedModel ClassIn-Dist AUROC (n=1,102)OOD AUROC (n=185)OOD Prec @ 0.62Unsafe Merges
JEV (System One)Probabilistic Choice0.8510.60550.0%15 / 30
Static Rule PolicyDeterministic Heuristic0.6020.55141.2%20 / 34
DeepSeek FlashGeneral LLM Prompt0.5850.51444.4%15 / 27
Note: In-distribution evaluated on 1,102 balanced cases (551 breaking, 551 control). Out-of-distribution evaluated on 185 cases (100 breaking, 85 control across 103 independent repos). Threshold evaluated at 0.62.

01 / Conceptual Framing: Separating Three Distinct Capabilities

In AI decision systems, three fundamental questions are routinely conflated: (1) Decision ranking quality: Does the model score safe updates above unsafe ones? (2) Probability calibration: Do the returned numeric probabilities retain their statistical meaning on unseen distributions? (3) Safe operational automation: Does a fixed confidence threshold enforce safe behavior in production without human intervention? An automated system can exhibit competitive discriminative ranking while catastrophically failing calibration and threshold safety.

02 / Experimental Setup & Hermetic Pre-Merge Isolation

To guarantee empirical rigor, all three systems (JEV, Static Rules, and DeepSeek Flash) evaluated identical pre-merge state: ecosystem, package manager, semantic version diff, PR title/body, manifest changes, release notes, and pre-merge CI status. Strictly zero post-merge telemetry (merge status, subsequent reverts, error logs) was accessible. Furthermore, models were not provided repository source code or API dependency graphs, isolating whether change metadata alone carries transferable decision signal.

// Pre-merge state payload structure evaluated by all systems (data/leakage-report.json)json
{
  "task": "DEPENDENCY_UPDATE_DECISION",
  "allowed_actions": ["AUTO_MERGE", "HOLD", "REQUIRE_REVIEW"],
  "pre_merge_state": {
    "repository": "isolated-repo-id",
    "ecosystem": "npm",
    "dependency": "package-name",
    "version_transition": { "from": "1.4.2", "to": "1.5.0", "semver_type": "minor" },
    "manifest_diff": "+ package-name@1.5.0\n- package-name@1.4.2",
    "ci_status_pre_merge": "SUCCESS",
    "has_release_notes": true
  },
  "leakage_boundary": {
    "post_merge_facts_included": false,
    "source_code_index_included": false
  }
}

03 / Empirical Findings: Discriminative Signal vs. Distributional Decay

On in-distribution evaluations (1,102 cases), JEV demonstrated strong discriminative capacity, achieving an AUROC of 0.851—decisively outperforming deterministic rules (0.602) and general-purpose LLM prompting (DeepSeek Flash at 0.585). However, when exposed to 185 out-of-distribution cases across 103 unseen repositories, JEV's AUROC dropped to 0.605 (95% CI: 0.525–0.685). While it remained the top-ranked system, the discriminative advantage over simple heuristics compressed drastically.

04 / Operational Failure: The Frozen Threshold Hazard

The critical finding of the benchmark is operational. In production automation, teams fix confidence thresholds (e.g. threshold = 0.62) to determine whether a pull request can bypass human review. Under distribution shift, this frozen threshold collapsed: out of 30 auto-merge executions, exactly 15 were breaking updates—a 50.0% precision rate equivalent to a coin toss. This empirical outcome establishes that static confidence thresholds cannot be safely frozen across shifting software codebases.

// Evaluation CLI reproduction commandbash
# Clone and reproduce the benchmark results locally
git clone https://github.com/scarif-labs/jev-software-decision-benchmark.git
cd jev-software-decision-benchmark
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python run-eval.py --dataset data/ood-185.json --model jev-latest

05 / Implications for Autonomous Software Engineering

For software engineers building agentic workflows and automated dependency systems: change metadata carries genuine predictive signal, but autonomous action requires dynamic calibration rather than static probability thresholds. Autonomous code operations must implement adaptive conformal prediction, rollback telemetry, and graduated human oversight when operating outside their training distribution.

Empirical Findings & Systems Takeaways
  • 01.Ranking signal does not equal operational safety: high AUROC (0.851) does not prevent unsafe actions if calibration decays under distribution shift.
  • 02.Frozen confidence thresholds are hazardous: an in-distribution threshold of 0.62 resulted in a 50.0% failure rate (15 unsafe merges) on out-of-distribution repos.
  • 03.Domain-specific decision primitives outperform general LLMs in ranking: JEV outperformed DeepSeek Flash across both in-distribution (0.851 vs 0.585) and OOD (0.605 vs 0.514).
  • 04.Software automation requires adaptive uncertainty estimation (such as conformal prediction) before autonomous write actions can be safely delegated.
Cite this empirical benchmark
@misc{scariflabs2026jevbenchmark,
  title={Independent Benchmark of JEV for Dependency-Update Automation Under Distribution Shift},
  author={Lawrance, Alen and PS, Harikrishnan and Prakash, Vishnu},
  year={2026},
  publisher={Scarif Labs Research},
  url={https://github.com/scarif-labs/jev-software-decision-benchmark},
  howpublished={\url{https://scariflabs.com/lab/jev-software-decision-benchmark}}
}

Applied in our production systems & services

Start a project

Interested in empirical AI evaluations or autonomous agent architecture?

We design, evaluate, and productionize intelligent systems and reliable software architectures.