{"language": "en", "source_path": "monthly/README.md", "markdown": "# Monthly Reports\n\n> **English** | [简体中文](../zh/monthly/README.md)\n\nMonthly reports explain what entered the knowledge base during one calendar month, what changed across the literature, and why those changes matter. They are repository changelogs with editorial synthesis, not publication-date archives.\n\n## How a month is assigned\n\nA work belongs to the month in which its card first reached `main`. Its card's separate `First appeared` stamp still records when the work itself first became publicly accessible.\n\nThe one-time historical archive is the exception: existing cards were backfilled into reports from January 2024 onward according to their public `First appeared` month. No historical reports are created for 2023 or earlier. Future reports return to the main-addition rule below, so a late-discovered older work remains visible as a backfill.\n\nThis distinction prevents two errors:\n\n- A newly released work is identified as a **New release** when its first-appearance month matches the report month.\n- An older work discovered later is identified as a **Backfill**. It still appears in the month when the repository learned about it, but the report does not describe it as newly published.\n\n## Report structure\n\nEvery report contains:\n\n1. **Month at a Glance** — counts and three to five conclusions worth carrying forward.\n2. **What Changed This Month** — three to six evidence-backed story lines that connect related works.\n3. **Selected Topic Developments** — only topics with a meaningful cluster, boundary change, or disagreement.\n4. **Selected Domain Developments** — only domains where the month's additions reveal a domain-specific constraint or shift.\n5. **Complete Monthly Index** — every card added that month, exactly once, with its first-appearance date, release/backfill status, topics, and domains.\n\nThe prose is selective; the index is exhaustive. A work may have several taxonomy labels, but the narrative should explain it in one primary location and cross-link the other relevant pages instead of repeating the same summary.\n\n## Archive Highlights\n\n<!-- MONTHLY_ARCHIVE_OVERVIEW_START -->\n- `32` reports currently cover `365` works from `2024-01` through `2026-08`.\n- The busiest report so far is [2026-08](./2026-08.md) with `40` works.\n- The archive is most consistently concentrated around `Scientific Agent Benchmarks, Trajectory Evaluation`.\n<!-- MONTHLY_ARCHIVE_OVERVIEW_END -->\n\n## Reports\n\nReports are listed newest first. Each row gives the size of the month, the main topic cluster, and one reason to open that report instead of treating the archive like a blind list.\n\n<!-- MONTHLY_REPORTS_START -->\n| Month | Works | Primary topics | Why revisit it |\n|---|---:|---|---|\n| [2026-08](./2026-08.md) | 40 | Scientific Agent Benchmarks, Trajectory Evaluation | August 2026 is the densest month in the archive so far, with 40 first appearances: professional workflows, experiment-efficient model discovery, trajectory diagnosis, and agent-improvement loops now appear as distinct evaluation problems rather than one benchmark trend. |\n| [2026-07](./2026-07.md) | 23 | Scientific Agent Benchmarks, Skill Hierarchy | July 2026 is smaller than June by raw count, but more coherent. |\n| [2026-06](./2026-06.md) | 36 | Scientific Agent Benchmarks, Skill Hierarchy | June 2026 is the month where this repository's expanded scope starts to look structurally justified rather than aspirational. |\n| [2026-05](./2026-05.md) | 28 | Scientific Agent Benchmarks, Trajectory Evaluation | May 2026 is a scale-up month just before the June breakout. |\n| [2026-04](./2026-04.md) | 18 | Scientific Agent Benchmarks, Trajectory Evaluation | April 2026 is a plan-and-failure month. |\n| [2026-03](./2026-03.md) | 15 | Scientific Agent Benchmarks, Trajectory Evaluation | March 2026 is a process-instrumentation month. |\n| [2026-02](./2026-02.md) | 13 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | February 2026 pushes evaluation deeper into institutional, biomedical, and safety-constrained settings. |\n| [2026-01](./2026-01.md) | 11 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | January 2026 is the month where exploratory science, long-document workflows, and step-level judging clearly meet. |\n| [2025-12](./2025-12.md) | 9 | Scientific Agent Benchmarks, Trajectory Evaluation | December 2025 feels like a maturity month for engineering-heavy scientific evaluation. |\n| [2025-11](./2025-11.md) | 6 | Scientific Agent Benchmarks, Resource-aware Evaluation | November 2025 is where cost and operational risk become hard constraints rather than side metrics. |\n| [2025-10](./2025-10.md) | 22 | Scientific Agent Benchmarks, Credit Assignment | October 2025 is a second breakout month, but this time the added breadth comes with more explicit diagnostic pressure. |\n| [2025-09](./2025-09.md) | 10 | Scientific Agent Benchmarks, Skill Hierarchy | September 2025 reads like an engineering-and-tooling month. |\n| [2025-08](./2025-08.md) | 5 | Scientific Agent Benchmarks, Benchmark Design, Validity & Contamination | August 2025 is a small month, but it is one of the sharpest months for benchmark design. |\n| [2025-07](./2025-07.md) | 10 | Scientific Agent Benchmarks, Survey | July 2025 is unusually reflective about who gets to define good performance. |\n| [2025-06](./2025-06.md) | 10 | Scientific Agent Benchmarks, Skill Hierarchy | June 2025 centers on professional-discipline evaluation under stronger execution and cost constraints. |\n| [2025-05](./2025-05.md) | 22 | Scientific Agent Benchmarks, Trajectory Evaluation | May 2025 is a breadth explosion with an early hint of the repository's later development-loop themes. |\n| [2025-04](./2025-04.md) | 8 | Scientific Agent Benchmarks, Evaluator Reliability & Validation | April 2025 is a verification-hardening month. |\n| [2025-03](./2025-03.md) | 6 | Scientific Agent Benchmarks, Survey | March 2025 is a framing month rather than a scale month. |\n| [2025-02](./2025-02.md) | 9 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | February 2025 is where scientific-agent environments and long-horizon evaluation start to meet. |\n| [2025-01](./2025-01.md) | 8 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | January 2025 is a breadth month that mixes frontier general exams with domain-heavy scientific tasks. |\n| [2024-12](./2024-12.md) | 6 | Scientific Agent Benchmarks, Credit Assignment | December 2024 reads like a transition from static benchmarks to live and process-aware evaluation. |\n| [2024-11](./2024-11.md) | 5 | Scientific Agent Benchmarks, Resource-aware Evaluation | November 2024 is small but methodologically mixed in a useful way. |\n| [2024-10](./2024-10.md) | 15 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | October 2024 is the first broad breakout month for agent evaluation across engineering, science, and embodied systems. |\n| [2024-09](./2024-09.md) | 4 | Scientific Agent Benchmarks, Skill Hierarchy | September 2024 is a decomposition month. |\n| [2024-08](./2024-08.md) | 4 | Scientific Agent Benchmarks, Trajectory Evaluation | August 2024 brings planning and execution closer to real operational settings. |\n| [2024-07](./2024-07.md) | 6 | Scientific Agent Benchmarks | July 2024 is the first month where domain-authentic scientific benchmarking clearly dominates. |\n| [2024-06](./2024-06.md) | 6 | Scientific Agent Benchmarks, Planning & Decision-Making Evaluation | June 2024 pushes evaluation toward external verification rather than plausible answers. |\n| [2024-05](./2024-05.md) | 4 | Scientific Agent Benchmarks | May 2024 is where real professional workflows begin to matter. |\n| [2024-04](./2024-04.md) | 2 | Scientific Agent Benchmarks, General Long-Horizon Agent Benchmarks | April 2024 is a two-work contrast that shows how much a headline score depends on the interaction surface. |\n| [2024-03](./2024-03.md) | 1 | Scientific Agent Benchmarks | March 2024 is small but important because BrainBench brings a full discipline into scope without pretending that general QA is enough. |\n| [2024-02](./2024-02.md) | 2 | General Long-Horizon Agent Benchmarks, Planning & Decision-Making Evaluation | February 2024 is the archive's first clean planning month. |\n| [2024-01](./2024-01.md) | 1 | Trajectory Evaluation, Skill Hierarchy | January 2024 is a one-paper launch month, but it already sets a durable standard for the archive. |\n<!-- MONTHLY_REPORTS_END -->\n"}
