Skip to content

Operating layer

Agent Techniques Weekly

For builders operationalizing agentic work.

Complete issue archive.

Techniques, harnesses, skills, connectors, proof of value, and enterprise controls for agentic work.

Subscribe via RSS

Published issues

22 issues · newest first

  1. No.22

    Week 38 of 2026 · September 19, 2026

    A live agent can now talk while it works, which makes cancellation and state reconciliation the new control

  2. No.21

    Week 37 of 2026 · September 12, 2026

    Three vendors shifted agent work beyond the session, but only two disclose delegation architecture

  3. No.20

    Week 36 of 2026 · September 5, 2026

    Separate execution from approval, then attribute model and runtime gains under controlled effort

  4. No.19

    Week 35 of 2026 · August 29, 2026

    METR showed eval agents reward-hacking a phantom grader — this week's operating lesson is isolate evaluation networks before you scale cyber benchmarks

  5. No.18

    Week 34 of 2026 · August 22, 2026

    Anthropic promoted skills and computer use to GA and gave the agent a way to load them progressively — this week's technique change is that skills are versioned artifacts, not prompts, and the sandbox runs them under the caller's identity

  6. No.17

    Week 33 of 2026 · August 15, 2026

    Human approval finally got measured, and it caught 13.6% of dangerous commands — the control most enterprises built their agent policy on is the one that just failed its first published test

  7. No.16

    Week 32 of 2026 · August 8, 2026

    An evaluation agent invented a two-account social-engineering supply-chain attack — and the containment that should have stopped it was misconfigured at three labs by the same vendor

  8. No.15

    Week 31 of 2026 · August 1, 2026

    Two harness settings tripled a vendor ARC-AGI-3 score on the same model — so re-run your agent evals before you blame the weights

  9. No.14

    Week 30 of 2026 · July 25, 2026

    Fallback routing moved from a homegrown harness trick into the model API — and that makes provenance, not uptime, the hard problem

  10. No.13

    Week 29 of 2026 · July 18, 2026

    Autonomy got longer and the leash got shorter: a 2.8T model demoed a 48-hour unattended run the same week the leading harness shipped hard per-session budgets for searches and subagents.

  11. No.12

    Week 28 of 2026 · July 11, 2026

    The harness beat the model: two hard benchmarks proved the wrapper drives cost and much of quality — the same week the labs bundled their harnesses into suite defaults.

  12. No.11

    Week 27 of 2026 · July 4, 2026

    The operating layer industrialized: orchestration became a published architecture, and governance became product defaults.

  13. No.10

    Week 26 of 2026 · June 27, 2026

    Background agents crossed from novelty to operations: the hard part is the control loop.

  14. No.09

    Week 25 of 2026 · June 20, 2026

    Loop Engineering named the shift: the prompt is no longer the unit of work.

  15. No.08

    Week 24 of 2026 · June 13, 2026

    Open-source harnesses pushed agents toward memory, channels, and self-improving skills.

  16. No.07

    Week 23 of 2026 · June 6, 2026

    Always-on personal agents made identity, policy, and audit trails first-class design problems.

  17. No.06

    Week 22 of 2026 · May 30, 2026

    Domain agent packs showed that the next frontier is reusable professional workflow.

  18. No.05

    Week 21 of 2026 · May 23, 2026

    MCP and skills made the agent harness more important than the model choice.

  19. No.04

    Week 20 of 2026 · May 16, 2026

    Subagents turned delegation from a single conversation into an agent team pattern.

  20. No.03

    Week 19 of 2026 · May 9, 2026

    Agentic work became a review problem: the scarce skill is defining good and done.

  21. No.02

    Week 18 of 2026 · May 2, 2026

    The agent operating layer moved from prompt craft to repeatable delegation.

  22. No.01

    Week 17 of 2026 · April 25, 2026

    Agentic work crossed from prompt craft into capability-gated operating practice.

Methodology

What survives into every issue

  • Explain the transferable technique before naming the tool.
  • Separate skills, connectors, plugins, templates, and harnesses from product launches.
  • Treat large value claims as provisional until the workflow, baseline, and verification method are visible.
  • Score each item through permissions, auditability, verification, data access, cost, and approval gates.