176-configuration harness ablation: rule-based elision before LLM summarization is the best context strategy, and making elided text recoverable barely helps
A Sep 17 arXiv study holds the execution loop fixed and varies planning, action space, and context management independently across four models and 176 configurations. Context management matters most under tight token budgets, the winning recipe is rule-based elision applied before LLM-based summarization, and making deleted content recoverable adds almost nothing. Planning acts as accuracy scaffolding for weaker models but only a cost reduction for stronger ones, and bash-only interfaces beat predefined tools for bash-proficient models, which is a concrete configuration rule rather than a general 'better harness' claim. This is a separate paper from the Berkeley HarnessTax dashboard covered on 09-17 and reaches component-level conclusions that dashboard did not.
↳ Follow the thread