Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Qwen2.5-3B-Instruct tested on HumanEval+ code generation benchmark.
Source findingRecuris improves on HumanEval+ benchmarks.
Source findingOckhamareto improves on HumanEval+ benchmarks.
Source findingThe paper tests post-training trajectories on HumanEval.
Source findingThe study evaluated coding agent reliability through forced revisions on HumanEval repairs.
Source findingIndustryCode benchmarks code generation in domains not covered by HumanEval.
Source findingEsoLang-Bench challenges validity of HumanEval benchmark.
Source findingpaddo.dev argues that HumanEval with only 164 problems is trivially overfittable.
Source findingTERMINATOR outperforms on HumanEval benchmark
Source findingRealBench shows significant performance drops compared to HumanEval benchmark.
Source findingDeepSeek V4 Lite reportedly achieves 90% HumanEval Pass@1
Source findingQwen2.5-3B-Instruct tested on HumanEval+ code generation benchmark.
Source finding