JevSpawn

Research preprint · October 2026

JevSpawn

Adaptive Agentic Inference through Compositional Action Spaces

Haoyang Su1,3Weiran Huang2,3

1 Fudan University2 Shanghai Jiao Tong University3 Shanghai Innovation Institute

An agent can explore many actions without writing each action from scratch. JevSpawn brings finite probabilistic prediction to multi-turn interaction through action spaces that are inferred from task context and adapted through feedback.

Task-balanced quality versus end-to-end latency, and quality accumulated over elapsed time, for JevSpawn and seven agent baselines.
Task quality and end-to-end latency across eight benchmark tasks.
8benchmark tasks
1,761evaluation instances
0additional training

From task meaning to finite actions

Natural-language rules are expressed as composable action fields. Finite probabilities drive parallel spawning, and returned observations guide continuation, revision, and recovery. Known alternatives share model computation instead of requiring a separate text generation for every branch.

Key results

Qwen3.8-27B · Four H100 GPUs

Maze success

0.96

40.91 s E2E latency

Grid success

0.95

40.58 s E2E latency

Highest task score

5 / 8

Against seven agent baselines

All task scores
Task score ↑. PPNL, Maze, and Grid report success rates. Other tasks use benchmark rewards or scores. Bold denotes the best result; underlining denotes the second best.
MethodPPNLMazeGridLightsOutRushHourSokoban2048Nullify
LATS0.8420.0000.0000.0100.0000.0000.0000.010
LLMCompiler0.9290.2000.4800.0800.1470.3505.0000.020
AgentPrune0.9560.2800.2800.3400.4530.7901.2400.030
HiAgent0.9300.4000.3200.0330.1720.1400.0000.090
FoldAgent0.4640.5200.7300.0200.2400.1050.0000.090
DyFlow0.8280.0400.0200.0400.0600.2000.1200.060
LatentMAS0.9020.3200.1200.0300.1130.0950.0000.050
TypeSafe Jev0.9550.9600.9300.6600.2900.250296.0000.170
JevSpawn0.9500.9600.9500.6100.3900.150305.1200.210

TypeSafe Jev uses the same JevSpawn architecture with API-based finite scoring.

End-to-end latency
E2E latency in seconds ↓. Bold denotes the best result; underlining denotes the second best.
MethodPPNLMazeGridLightsOutRushHourSokoban2048Nullify
LATS131.34300.00300.00299.01298.57300.00300.00298.18
LLMCompiler13.61235.32117.1699.26223.42123.9933.51172.85
AgentPrune24.2347.9179.57233.19115.48142.53202.5742.90
HiAgent16.25209.06206.05254.5420.32150.76300.0098.18
FoldAgent43.55191.6961.4484.1722.7478.32300.0748.10
DyFlow125.47285.43269.57294.18289.03280.57287.35287.37
LatentMAS48.70269.86281.61292.64150.25269.05289.03185.05
TypeSafe Jev43.9663.9878.39153.95187.45214.76249.05189.41
JevSpawn21.8840.9140.58111.0589.20122.49169.94104.21

E2E latency includes failed tasks and timeouts. API communication is included for TypeSafe Jev.

Watch the interaction unfold

Follow eight paper case studies through action declarations, spawned branches, observations, and answers.

The video replays recorded traces. Run the demo locally to explore each turn.

bash demo/run.sh

Run the demo ↗

Citation

Download .bib
@misc{su2026jevspawnadaptiveagenticinference,
      title={JevSpawn: Adaptive Agentic Inference through Compositional Action Spaces},
      author={Haoyang Su and Weiran Huang},
      year={2026},
      eprint={2610.00437},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2610.00437},
}