ACCEPTED TO NEURIPS 2026

ACTION ADVANTAGE ASSIGNMENT

Learning CLI Agents
with Structured Action Credit
under Selective Observation

Haoyang Su1,3 · Ying Wen2,3

1 Fudan University2 Shanghai Jiao Tong University3 Shanghai Innovation Institute

Use the structure of executable actions to learn from multi-turn interaction.

What should an agent see?
Which actions deserve credit?

CLI agents navigate evolving filesystems, execute commands, and learn from delayed task feedback. Two bottlenecks are closely connected: selecting relevant workspace evidence under partial observation, and assigning returns across a multi-turn trajectory.

We introduce A3, an agentic reinforcement learning method that builds advantages from shell command structure, and σ-Reveal, an inference-time harness that selects budgeted initial workspace context. ShellOps provides verifiable tasks for studying these capabilities.

24.6%

ShellOps Hybrid exact match

A3 + σ-Reveal · best compared non-A3 baseline: 11.3%
55.7%

ShellOps Pass@5

A3 + σ-Reveal · best compared non-A3 baseline: 28.6%
0.023s

Advantage computation

Reported training setting · GSPO: 0.027 s

With Qwen3-14B, A3 + σ-Reveal more than doubles the best compared baseline's ShellOps Hybrid exact match. A3 constructs advantages from sampled trajectories without an additional critic or an external LLM judge.

Action structure becomes a learning signal.

A3 represents shell commands through abstract syntax trees and compares sequences of actions.

Episode feedback, action residuals, and branch returns jointly determine the advantage at each turn.

A3 method diagram: AST action comparison, Sigma-Reveal workspace selection, and episode, turn, and tree advantages.
Method overview. σ-Reveal supplies initial workspace context; A3 combines episode feedback with return residuals computed for similar action sequences and credit assigned across abstract histories.
01

Episode feedback

Normalize returns across sibling rollouts for the same task to retain a task-level learning signal.

02

Action sub-chain residuals

Compare returns within groups of structurally similar action sequences at the same turn, across multiple sequence lengths.

03

Abstract-history branches

Compare branch returns under similar action histories and accumulate the resulting margins along the trajectory.

σ-REVEAL

Select workspace context before the first action.

σ-Reveal uses filenames mentioned in the task, directory depth, and file extensions to score entries in the workspace. It selects a set that fits the context budget and includes the parent directories of selected files.

The resulting view is added to the initial prompt. The agent then explores the workspace through shell commands and their execution results.

Stronger learning on composite workspace tasks.

We compare methods using Qwen3-14B on tasks that require correct answers, file edits, or both.

The table reports exact match on each of the three ShellOps task types.

ShellOps exact match (%) · Qwen3-14B
MethodStringFilesHybrid
ReACT26.07.17.2
LATS27.510.99.1
rStar18.58.77.0
GSPO25.510.911.3
GiGPO24.011.310.1
HGPO23.511.99.5
RetroAgent19.58.29.7
A346.526.521.9
A3 + σ-Reveal48.525.724.6

σ-Reveal is applied at inference time. See the paper for the full six-stream evaluation.

Transfer to ShellOps-Pro

The same trained policy is evaluated on 150 harder tasks. With σ-Reveal, A3 reaches 29.9%, 30.5%, and 31.1% macro exact match at 6, 8, and 10 turns.

At 10 turns, the reported Qwen3-235B-A22B baseline scores 28.3%; GLM-5.1 and Kimi-K2.6 remain ahead at 40.7% and 53.3%. These comparisons use the paper's shared shell evaluation setting.

Structure matters

Removing both the turn and tree channels reduces mixed-benchmark Hybrid exact match from 21.9% to 10.9%. Component effects vary across task types.

AST similarity is a structural proxy for action intent. It does not establish semantic equivalence or identify the causal contribution of an action.

Training curves comparing success, answer reward, policy entropy, and PPO surrogate KL.
Training dynamics under matched data and rollout settings. A3 maintains controlled updates while success and reward improve.
Advantage composition and computational efficiency +
Advantage component magnitudes and ShellOps accuracy versus advantage computation cost.
Component shares describe weighted signal magnitudes. Costs refer to advantage computation.

Evaluate what happens in the filesystem.

A verifiable task suite for information extraction and file editing through executable shell interaction.

1,624standard corpus tasks

150ShellOps-Pro tasks

4,063files across Pro workspaces

ShellOps task structure, lookup, aggregate, edit and mixed task categories, and the verifiable agent-environment loop.
Tasks pair an instruction with an initial workspace and verifiable output or target file state. The paper uses 714 in-distribution ShellOps tasks for training and evaluation; ShellOps-Pro provides harder out-of-distribution workspaces.
Explore ShellOps on Hugging Face ↗

Build on this work.

If you use A3, σ-Reveal, or ShellOps in your research, please cite the paper.

@misc{su2026learningcliagentsstructured,
      title={Learning CLI Agents with Structured Action Credit under Selective Observation},
      author={Haoyang Su and Ying Wen},
      year={2026},
      eprint={2605.08013},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2605.08013},
}

Download citation.bib ↓