Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
CogEvol training uses GRPO, a hybrid rule-plus-VLM reward method.
Source findingSAPO beats GRPO by 12.1 percentage points on agent benchmarks.
Source findingAmazon Nova Lite 2.0 was trained using GRPO with four-component rewards.
Source findingTraVEL applies Group Relative Policy Optimization with ego-trajectory similarity reward.
Source findingMemHarness is trained end-to-end with GRPO.
Source findingGRPO is the algorithm behind DeepSeek R1
Source findingUnsloth released long-context GRPO with batching algorithms
Source findingDeepSeek-R1 uses GRPO algorithm for optimization.
Source findingTRIAGE addresses GRPO's flaw of uniform advantage that punishes useful exploration.
Source findingGradient trains agents using GRPO.
Source findingCoPES recovers 92% of GRPO's validation-accuracy gains with one-eighth the GPU memory.
Source findingGRPO system uses mixed-integer programming formulations as candidate solutions
Source finding