Episode feedback
Normalize returns across sibling rollouts for the same task to retain a task-level learning signal.
ACTION ADVANTAGE ASSIGNMENT
1 Fudan University2 Shanghai Jiao Tong University3 Shanghai Innovation Institute
Use the structure of executable actions to learn from multi-turn interaction.
CLI agents navigate evolving filesystems, execute commands, and learn from delayed task feedback. Two bottlenecks are closely connected: selecting relevant workspace evidence under partial observation, and assigning returns across a multi-turn trajectory.
We introduce A3, an agentic reinforcement learning method that builds advantages from shell command structure, and σ-Reveal, an inference-time harness that selects budgeted initial workspace context. ShellOps provides verifiable tasks for studying these capabilities.
ShellOps Hybrid exact match
A3 + σ-Reveal · best compared non-A3 baseline: 11.3%ShellOps Pass@5
A3 + σ-Reveal · best compared non-A3 baseline: 28.6%Advantage computation
Reported training setting · GSPO: 0.027 sWith Qwen3-14B, A3 + σ-Reveal more than doubles the best compared baseline's ShellOps Hybrid exact match. A3 constructs advantages from sampled trajectories without an additional critic or an external LLM judge.
A3 represents shell commands through abstract syntax trees and compares sequences of actions.
Episode feedback, action residuals, and branch returns jointly determine the advantage at each turn.

Normalize returns across sibling rollouts for the same task to retain a task-level learning signal.
Compare returns within groups of structurally similar action sequences at the same turn, across multiple sequence lengths.
Compare branch returns under similar action histories and accumulate the resulting margins along the trajectory.
σ-Reveal uses filenames mentioned in the task, directory depth, and file extensions to score entries in the workspace. It selects a set that fits the context budget and includes the parent directories of selected files.
The resulting view is added to the initial prompt. The agent then explores the workspace through shell commands and their execution results.
We compare methods using Qwen3-14B on tasks that require correct answers, file edits, or both.
The table reports exact match on each of the three ShellOps task types.
| Method | String | Files | Hybrid |
|---|---|---|---|
| ReACT | 26.0 | 7.1 | 7.2 |
| LATS | 27.5 | 10.9 | 9.1 |
| rStar | 18.5 | 8.7 | 7.0 |
| GSPO | 25.5 | 10.9 | 11.3 |
| GiGPO | 24.0 | 11.3 | 10.1 |
| HGPO | 23.5 | 11.9 | 9.5 |
| RetroAgent | 19.5 | 8.2 | 9.7 |
| A3 | 46.5 | 26.5 | 21.9 |
| A3 + σ-Reveal | 48.5 | 25.7 | 24.6 |
σ-Reveal is applied at inference time. See the paper for the full six-stream evaluation.
The same trained policy is evaluated on 150 harder tasks. With σ-Reveal, A3 reaches 29.9%, 30.5%, and 31.1% macro exact match at 6, 8, and 10 turns.
At 10 turns, the reported Qwen3-235B-A22B baseline scores 28.3%; GLM-5.1 and Kimi-K2.6 remain ahead at 40.7% and 53.3%. These comparisons use the paper's shared shell evaluation setting.
Removing both the turn and tree channels reduces mixed-benchmark Hybrid exact match from 21.9% to 10.9%. Component effects vary across task types.
AST similarity is a structural proxy for action intent. It does not establish semantic equivalence or identify the causal contribution of an action.


A verifiable task suite for information extraction and file editing through executable shell interaction.
1,624standard corpus tasks
150ShellOps-Pro tasks
4,063files across Pro workspaces

If you use A3, σ-Reveal, or ShellOps in your research, please cite the paper.
@misc{su2026learningcliagentsstructured,
title={Learning CLI Agents with Structured Action Credit under Selective Observation},
author={Haoyang Su and Ying Wen},
year={2026},
eprint={2605.08013},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.08013},
}Download citation.bib ↓