AI Coding Tools
AI coding assistants, IDEs, and developer workflows
Living document tracking AI coding assistants, IDEs, and developer workflows. Newest entries appear at the top.
Key Areas to Watch
- IDE integrations (Cursor, GitHub Copilot, Claude Code, Windsurf)
- Code generation and completion models
- Automated testing and debugging
- Repository-scale understanding
- Developer experience and productivity studies
2026-W39
- Claude Code now reads AGENTS.md — Fallback support in version 2.1.277 improves project-instruction portability across coding-agent tools. [HN]
- Replace PRs with Delta — now in public beta — Zed is testing a change-review workflow that may better fit agent-produced work than conventional pull requests. [Zed]
- Tell agents the why, not just the how — Intent-rich instructions let stronger coding agents recover from assumptions without overfitting to brittle implementation steps. [Sean Goedecke]
- CodeMidas — Generates executable RL tasks directly from implemented functionality across open-source codebases. [arXiv]
2026-W38
- Cognition helps Devin test its own work with GPT-6 Astra — Concrete pattern for agents validating other agents’ software work. [OpenAI]
- Slow developer experience will bottleneck fast models — Developer tooling around the model remains the bottleneck even as model latency improves. [Sean Goedecke]
- I-have-ADHD — Small prompt/skill that forces coding agents to stop burying the answer, illustrating how workflow rules shape agent usability. [HN]
- Le Chat custom MCP connectors — MCP connector and memory workflows continue moving into mainstream agent products. [Mistral]
2026-W37
- Give Your Coding Agents a Memory You Own — Owned memory for coding agents is becoming part of the team workflow stack. [Hugging Face]
- Your Agent Speaks MCP. Give It a Computer. — Agent runtimes are moving toward real computers with inspectable state and MCP-facing tools. [Fly.io]
- Using Blender with coding agents on macOS — Practical example of coding agents operating local creative software through available interfaces. [Simon Willison]
2026-W36
- GPT-5.6 in Kiro — OpenAI framed GPT-5.6 around developer price-performance in a coding environment. [OpenAI]
- Admin plugin for ChatGPT Work and Codex — Team administration is becoming part of operating AI coding systems. [OpenAI]
- Coding expertise is going to collapse from AI reliance — Developer-skill critique worth tracking alongside agent adoption. [HN]
2026-W32
- Stacked PRs are now live on GitHub — Native stacked changes make agent-generated work easier to split and review. [HN]
- LLMs reward expertise — Better developers can extract more from AI systems because they steer, inspect, and recover better. [HN/RSS]
- Codex Security — Security-focused Codex workflows keep vulnerability finding and repair in the coding-agent lane. [HN]
2026-W28
- Claude Code is steganographically marking requests — Hidden request metadata turned coding-agent CLIs into a transparency and trust-boundary discussion. [HN]
- Godot will no longer accept AI-authored code contributions — Maintainers are formalizing provenance rules for AI-generated patches. [HN]
- ScarfBench — Enterprise Java migration became a benchmarkable agent workload. [Hugging Face]
2026-W27
- Codex-maxxing for long-running work — OpenAI documented Codex workflows for sustained engineering projects. [OpenAI]
- Codex logging bug may write TBs to local SSDs — A local logging bug highlighted why agent CLIs need operational guardrails. [HN]
- Reusable Codex permissions — OpenAI developer posts described reusable, inheritable permission sets as a replacement for coarse sandbox modes. [x.com]
2026-W26
- Daybreak and Codex Security — OpenAI positions Codex as a security workflow for finding, validating, and patching vulnerabilities. [OpenAI]
- Patch the Planet — AI-assisted open-source security initiative with maintainer and expert-review loops. [OpenAI]
- Mistral Code — Coding assistant product bundling models, IDE support, local deployment, and enterprise controls. [Mistral]
- Devstral 2 and Mistral Vibe CLI — Coding models and a CLI surface for agentic development workflows. [Mistral]
- Codex-maxxing for long-running work — Practical example of Codex as a persistent engineering context manager. [OpenAI]
- Codex logging bug may write TBs to local SSDs — Operational bug showing why coding-agent CLIs need quotas, log rotation, and predictable local-resource use. [HN]
2026-W25
- How engineers at Nextdoor use Codex to build without limits — Codex appears in another concrete engineering-team workflow case study. [OpenAI]
- What Codex unlocks for Notion — Notion’s case study frames coding agents as a way to spread implementation work across product teams. [OpenAI]
- How an astrophysicist uses Codex to help simulate black holes — Codex use in scientific computing broadens the coding-agent story beyond web applications. [OpenAI]
- Homebrew 6.0.0 — Package-manager reliability remains part of the substrate coding agents depend on for setup and repair loops. [HN]
- macOS Container Machines — macOS containerization drew developer attention as a possible local sandbox surface for repeatable agent work. [HN]
2026-W24
- Codex for every role, tool, and workflow — Codex added plugins, sites, and annotations that move coding-agent UX into more team surfaces. [OpenAI]
- How Wasmer used Codex to build a Node.js runtime for the edge — Case study of agent-assisted runtime engineering with Codex and GPT-5.5. [OpenAI]
- Remote agents in Vibe. Powered by Mistral Medium 3.5 — Mistral added async remote coding agents that complete work outside the local editor session. [Mistral]
- Uber’s $1,500/month AI limit is a useful signal for AI tool pricing — AI coding-tool quotas and caps are becoming design constraints for agent loops and team workflows. [Simon Willison]
- What GitHub Copilot’s Usage-Based Billing Means for Zed Users — Zed clarified the practical impact of Copilot Chat metering for editor users. [Zed]
- Do Coding Agents Deceive Us? — Randomized capped tests target shortcut exploitation in coding-agent evaluations. [arXiv]
- SWE-Explore — Benchmark isolates how coding agents explore repositories before patching. [arXiv]
2026-W23
- Using AI to write better code more slowly — Nolan Lawson’s widely shared essay argues for using coding agents to improve deliberation and review quality rather than only speed. [HN]
- What GitHub Copilot’s Usage-Based Billing Means for Zed Users — Zed analyzed how Copilot’s billing shift changes the economics of editor-integrated AI assistance. [Zed]
- Build agents, not pipelines — Sean Goedecke makes the case for flexible agent loops over brittle static automation in changing developer workflows. [RSS]
- DeepSWE benchmark discussion — Developer attention moved toward agentic coding benchmarks that better match real software work than saturated SWE-Bench-style evals. [x.com]
- Please Do Not Vibe Fuck Up This Software — Rsync maintainers’ issue captured growing OSS pushback against low-quality AI-generated contributions to critical projects. [HN]
2026-W22
- Devstral — Mistral and All Hands AI release an Apache 2.0 agentic LLM for software-engineering tasks. [Mistral]
- OpenAI Codex at Virgin Atlantic and Ramp code review — Codex gets more production case studies in shipping and review workflows. [OpenAI]
- OpenAI and Dell bring Codex to hybrid/on-prem enterprise environments — On-prem coding agents become a clearer enterprise deployment lane. [OpenAI]
- DeepSeek Reasonix — DeepSeek-native coding agent draws strong developer interest. [HN]
2026-W20
- xAI Grok Build — A terminal-based agentic coding CLI for professional software engineering, in early beta behind a new $299/mo Grok SuperHeavy tier; every frontier lab now ships a coding CLI. [Engadget]
- OpenAI Codex comes to your phone — Codex lands in the ChatGPT mobile app in preview, letting developers view live environments, review outputs, and approve commands from a phone. [TechCrunch]
- Claude for Small Business — Anthropic extends its coding/enterprise push down-market; separately, Claude Code weekly limits rose 50% through July 13. [Anthropic / Reddit]
- Semble — code search for agents — Claims 98% fewer tokens than grep for agent context-gathering, targeting the token-cost bottleneck in agentic coding loops. [HN]
- Code as Agent Harness — Argues for treating generated code itself as the agent’s harness/scaffolding, formalizing a pattern now common in production agentic systems. [arXiv]
- Developer revolt over Claude Code third-party limits — As first-party Claude Code limits rose 50%, developers said Anthropic cut effective usage ~25x for third-party harnesses (T3 Code, Conductor, Zed,
claude -pin CI); Theo Browne canceled publicly and pledged $10/cancellation to open source. A live tension between first-party generosity and ecosystem lock-down. [x.com]
2026-W19
- Code with Claude — Anthropic’s first developer conference — Simon Willison live-blogged the full event, covering Claude Code workflows, extended context patterns, and agentic loop design. [Anthropic / Simon Willison]
- Vibe coding and agentic engineering are getting closer than I’d like — Simon Willison distinguishes vibe coding (fast HTML throwaway) from agentic engineering (artifact-producing loops); required reading for teams choosing a Claude Code strategy. [Simon Willison / HN]
- Karpathy: ask your LLM to structure responses as HTML — 15k-like tweet that sparked a wave of HTML-output experiments; directly inspired the Code with Claude conference session on HTML-first agent outputs. [x.com / @karpathy]
- AlphaEvolve: Gemini-powered coding agent scaling impact across fields — DeepMind’s detailed write-up on the evolutionary coding agent; already recovering matrix-multiply improvements and FLOP-efficient scheduling in production. [DeepMind Blog]
- Zed Zeta2.1: 3x fewer tokens, 50ms faster — Zed’s inline completion model cuts token count 3x and latency 50ms from Zeta2.0. [Zed Blog]
- Zed for Business — Team admin, SSO, and audit-log tiers ship; Zed makes its enterprise pitch official. [Zed Blog]
- Peekaboo 3.0: action-first macOS computer use — Steipete ships Peekaboo 3.0 with unified screen capture, computer-use agent, and new orchestration layer. (6,292 likes) [x.com / @steipete]
- AI makes weak engineers less harmful — Sean Goedecke argues AI raises the floor for low-skill contributions more than it raises the ceiling; practical framing for teams managing mixed-skill engineers. [Sean Goedecke]
- LLMs corrupt your documents when you delegate — Empirical study finding that LLM-delegated document edits introduce subtle factual and formatting corruption; important for agentic document workflows. [arXiv / HN]
2026-W18
- Remote agents in Vibe + Mistral Medium 3.5 — Cloud-async coding agents that complete sessions and PR back; Devstral and Magistral fold into a single Medium 3.5 path. [Mistral]
- GitHub Copilot is moving to usage-based billing — Individual plans convert on April 27; ends the unlimited-plan era. 767 HN points. [HN]
- VS Code inserting ‘Co-Authored-by Copilot’ into commits regardless of usage — 1,507-point HN thread on VS Code attributing commits even when Copilot wasn’t invoked. [HN]
- Claude Code refuses requests or charges extra if your commits mention “OpenClaw” — 1,342-point HN flag on string-matching billing behavior in Claude Code. [HN / x.com]
- HERMES.md in commit messages causes requests to route to extra usage billing — Anthropic GitHub issue for the same class of edge case. 1,250 HN points. [HN]
- DeepClaude — Claude Code agent loop with DeepSeek V4 Pro — Open-source agent harness pairing Claude Code’s planner with DeepSeek V4-Pro’s 1M context. 673 HN points. [HN]
- The Zig project’s anti-AI contribution policy — Simon Willison breakdown of Zig’s stated reasoning. 681 HN points. [HN]
- Codex CLI 0.128.0 adds /goal — OpenAI’s Codex CLI adds a goal-tracking primitive useful for long-horizon agent runs. [Simon Willison]
- llm 0.32a0 is a major backwards-compatible refactor — Simon Willison’s CLI/library hits a major architectural milestone. [Simon Willison]
- AlphaEvolve: How our Gemini-powered coding agent is scaling impact across fields — DeepMind details a Gemini-based coding agent applied across math, infra, and chip design. [DeepMind]
- Symphony orchestration spec, Codex edition — Codex-side reference implementation for OpenAI’s open-source orchestration spec. [OpenAI]
- Who owns the code Claude Code wrote? — 555-point HN legal analysis on AI-generated code ownership. [HN]
- I gave Claude Code a $0.02/call coworker and stopped hitting Pro limits — 1,754-point r/ClaudeAI workflow pattern using a cheap secondary model to absorb status traffic. [Reddit]
- Prompt Injection experience — my first time ever — 1,323-point r/ClaudeAI incident write-up. [Reddit]
- I accidentally burned ~$6,000 of Claude usage overnight with one command — Cautionary tale on unbounded sub-agent loops. 1,262 points. [Reddit]
2026-W17
- An update on recent Claude Code quality reports — Anthropic’s postmortem: three harness bugs, not models, drove two months of complaints. 939 HN points. [HN / Anthropic]
- Is Claude Code going to cost $100/month? Probably not — Anthropic silently removed Claude Code from the Pro plan, then restored it after backlash. [Simon Willison]
- I cancelled Claude: Token issues, declining quality, and poor support — 965-point HN frustration thread. [HN]
- Anthropic admits to making hosted models more stupid — 1,252-point r/LocalLLaMA framing of the postmortem as a pitch for open weights. [Reddit]
- Scaling Codex to enterprises worldwide — Codex Labs launches with Accenture, PwC, and Infosys; 4M weekly active users disclosed. [OpenAI]
- Top 10 uses for Codex at work — OpenAI’s playbook for non-engineering Codex deployments. [OpenAI]
- Speeding up agentic workflows with WebSockets in the Responses API — Connection-scoped caching dropped Codex’s API overhead substantially. [OpenAI]
- Introducing workspace agents in ChatGPT — Codex-powered ChatGPT agents for team workflows. [OpenAI]
- Introducing Parallel Agents in Zed — Run multiple coding agents at once in the same window. [Zed]
- Symphony: open-source spec for agent orchestration — OpenAI publishes an interoperability spec for multi-agent orchestration; positioning play against agent-platform lock-in. [OpenAI]
- clawsweeper: 50 codex agents triaging issues and PRs — Steinberger’s maintainer-tier agentic janitor closed 10K issues and 5K PRs in a single week. [x.com]
- Codex App update — full browser use, global dictation, non-dev mode, auto-review — Same-hour drop alongside GPT-5.5 launch. [x.com]
- Changes to GitHub Copilot Individual plans — Pricing and feature changes for Copilot Individual; 540 HN points. [HN]
- Over-editing: a model modifying code beyond what is necessary — 422-point HN essay on the most common Claude/Cursor failure mode. [HN]
- Anthropic says OpenClaw-style Claude CLI usage is allowed again — Reverses last week’s third-party-harness blocks. 510 HN points. [HN]
- An AI agent deleted our production database — 802-point HN incident postmortem; agent confession appended. [HN]
- PSA: HERMES.md in your git history routes Claude Code billing to extra usage — 1,372-point r/ClaudeAI warning on a string-matching billing edge case. [Reddit]
- SpaceX agreement to acquire Cursor for $60B — Or pay $10B for “our work together.” 822 HN points. [HN]
- llm 0.31 ships GPT-5.5 support and a verbosity option — Simon Willison’s CLI gets a verbosity knob and the new model. [Simon Willison]
- llm-openai-via-codex 0.1a0 — Plugin that hijacks Codex CLI credentials to call GPT-5.5 from
llm. [Simon Willison] - Software engineering may no longer be a lifetime career — Sean Goedecke on whether AI-native coding flattens skill acquisition for new engineers. [Sean Goedecke]
- Claude Code cheat sheet after 6 months of daily use — Practical workflow tips compiled by an active practitioner. [Reddit]
2026-W16
- Claude Design — Anthropic Labs’ design tool; Figma closed down 4.26% on launch day. 1,224 HN points. [HN / Anthropic]
- Claude Code Routines — Formalized long-running workflows inside Claude Code; 718 HN points. [HN]
- Codex for (almost) everything — OpenAI expands Codex with computer use, in-app browser, image generation, memory, and plugins. [OpenAI]
- The next evolution of the Agents SDK — Native sandbox execution and model-native harness for long-running agents. [OpenAI]
- Cloudflare’s AI Platform: an inference layer designed for agents — Cloudflare ships an agent-first inference stack. [Cloudflare]
- Changes in the system prompt between Claude Opus 4.6 and 4.7 — Simon Willison diff of the published system prompts. [Simon Willison]
- Claude system prompts as a git timeline — Research repo turning Anthropic’s prompt archive into a git history. [Simon Willison]
- Claude Code workflow tips after 6 months of daily use — 972-point senior-dev practical writeup. [Reddit]
- Opus 4.7 is an over-engineering master — Cursor users note Opus 4.7 defaults toward heavier solutions than 4.6. [Reddit]
- Headless everything for personal AI — Simon Willison on Matt Webb’s argument that headless APIs will beat GUI-scraping agents. [Simon Willison]
- GitHub Stacked PRs — First-party stacked-PR primitive on GitHub; 898 HN points. [HN]
- jj – the CLI for Jujutsu — Ongoing Jujutsu traction in agentic-coding circles; 548 HN points. [HN]
- Smol machines: subsecond coldstart portable VMs — Sandbox primitive for emerging agent runtimes. [HN]
- Peter Steinberger — Summarize 0.13 ships, now official Homebrew formula — GitHub Copilot as a model backend, local video-slide extraction. [x.com]
- Steipete: Anthropic blocks first-party harness use for third-party wrappers —
claude -p --append-system-prompt400s for OpenClaw-style clients. [x.com]
2026-W15
- Introducing Zed’s Agent Metrics — Zed ships in-editor agent usage and quality metrics, targeting the same telemetry gap Cursor and Claude Code have left open. [Zed Blog]
- How We Developed Zeta2 — Zed’s engineering walkthrough of training its next-gen in-editor completion model. [Zed Blog]
- Reallocating $100/Month Claude Code Spend to Zed and OpenRouter — Public migration post that captured the week’s Claude Code discontent. [HN]
- The golden age is over — 3,760-point r/ClaudeAI thread; community consensus that Claude Code quality has regressed. [Reddit]
- Anthropic: Stop shipping. Seriously. — 3,107-point follow-up documenting release-cycle regressions. [Reddit]
- Pro Max 5x quota exhausted in 1.5 hours despite moderate usage — Official GitHub issue documenting quota-burn regressions on the Pro Max 5x tier. [GitHub]
- Anthropic downgraded cache TTL on March 6th — Cache-TTL change identified as likely driver of the quality + quota complaints. [GitHub]
- AI assistance when contributing to the Linux kernel — Official Linux kernel documentation on AI coding assistants for contributors. [GitHub]
- Codex for (almost) everything — OpenAI repositions Codex as a general-purpose coding agent. [OpenAI]
- The next evolution of the Agents SDK — Updated tool-use and orchestration primitives for OpenAI’s Agents SDK. [OpenAI]
2026-W14
- Claude Code’s source code has been leaked via a map file in their NPM registry — Original discovery thread exposing Claude Code internals including fake tool stubs, frustration-detection regexes, and an undercover mode; 2,095 HN points. [HN]
- The Claude Code Source Leak: fake tools, frustration regexes, undercover mode — Deep-dive analysis of the leaked NPM source map, giving developers unprecedented visibility into how a production agentic CLI is assembled; 1,376 HN points. [HN]
- Claude Code Unpacked: A visual guide — Visual breakdown of Claude Code’s internal architecture from the leak; 1,128 HN points. [HN]
- Tell HN: Anthropic no longer allowing Claude Code subscriptions to use OpenClaw — Anthropic locks down OpenClaw proxy access for Claude Code subscribers; 1,099 HN points, 829 comments. [HN]
- Claude Code Found a Linux Vulnerability Hidden for 23 Years — Claude Code surfaces a 23-year-old Linux bug; 433 HN points. [HN]
- Universal Claude.md – cut Claude output tokens — Community reusable CLAUDE.md pattern for reducing token output; 471 HN points. [GitHub]
- Lemonade by AMD: a fast and open source local LLM server using GPU and NPU — AMD’s open local inference server with GPU and NPU routing; 572 HN points. [HN]
- Running Gemma 4 locally with LM Studio’s new headless CLI and Claude Code — Practical walkthrough combining LM Studio headless CLI and Claude Code for local Gemma 4; 406 HN points. [HN]
- Codex now offers more flexible pricing for teams — OpenAI adds PAYG pricing for ChatGPT Business and Enterprise Codex users. [OpenAI]
- Programming (with AI agents) as theory building — Sean Goedecke applies Naur’s theory-building view to AI-agent-assisted development. [Sean Goedecke]
2026-W13
- Claude Code Cheat Sheet — Community-built quick-reference for Claude Code power users; 699 HN points, 188 comments. [HN]
- Anatomy of the .claude/ folder — Practical deep-dive into the Claude Code project configuration directory; 627 HN points, 265 comments. [HN]
- How I’m Productive with Claude Code — Workflow patterns from a practitioner using Claude Code in production; 281 HN points. [HN]
- Schedule tasks on the web — Claude Code docs — New Claude Code web-scheduled tasks feature; 298 HN points. [Anthropic]
- Zeta2: Zed’s edit prediction model is 30% better than Zeta1 — Zed ships a rebuilt inline edit prediction model trained from scratch with full data pipeline details. [Zed Blog]
- Claude Code runs git reset –hard origin/main every 10 minutes — Reported disruptive bug in Claude Code’s autonomous loop that destroys local state; 251 HN points. [GitHub]
- 90% of Claude-linked output going to GitHub repos with <2 stars — Aggregate stat suggesting most Claude Code usage produces private or near-private code; 337 HN points. [HN]
- We rewrote JSONata with AI in a day, saved $500k/year — Concrete ROI story for AI-assisted library rewrite; 276 HN points. [HN]
2026-W12
- OpenCode – Open source AI coding agent — Open-source CLI coding agent alternative to Claude Code and Cursor; 1,274 HN points, 619 comments on launch day. [opencode.ai]
- Mistral AI Releases Forge — Mistral’s new coding agent product; 733 HN points, 193 comments. [Mistral AI]
- Leanstral: Open-source agent for trustworthy coding and formal proof engineering — Mistral’s Lean 4 formal proof agent; 783 HN points. [Mistral AI]
- OpenAI to acquire Astral — OpenAI buys the team behind the Ruff Python linter and uv package manager to accelerate Codex and Python developer tooling. [OpenAI]
- Push events into a running session with channels — Claude Code gains session channels for pushing external events into live coding sessions; 400 HN points. [Anthropic]
- Get Shit Done: A meta-prompting, context engineering and spec-driven dev system — Opinionated open-source spec-driven dev system for LLM coding workflows; 473 HN points. [GitHub]
- Why Codex Security Doesn’t Include a SAST Report — OpenAI explains the AI-driven constraint reasoning approach replacing traditional SAST in Codex Security. [OpenAI]
- Show HN: Sub-millisecond VM sandboxes using CoW memory forking — Sub-ms VM sandboxing via CoW forking, directly relevant to fast agent code execution; 311 HN points. [GitHub]
2026-W11
- After outages, Amazon to make senior engineers sign off on AI-assisted changes — Amazon policy requiring senior-engineer approval on all AI-coding changes, following a string of production outages attributed to AI tools; 659 HN points, 485 comments. [Ars Technica]
- 1M context is now generally available for Opus 4.6 and Sonnet 4.6 — Anthropic removes limited-access friction for 1M-token context windows, unblocking long-context agent and retrieval pipelines; 1,220 HN points. [Anthropic]
- From model to agent: Equipping the Responses API with a computer environment — How OpenAI wired shell tools and hosted containers into the Responses API for stateful agent execution. [OpenAI]
- Designing AI agents to resist prompt injection — Technical post on constraining risky agent actions and guarding sensitive data in ChatGPT workflows. [OpenAI]
- Improving instruction hierarchy in frontier LLMs — IH-Challenge training to make models respect trusted-instruction priority and resist injection. [OpenAI]
- Unfortunately, Sprites Now Speak MCP — Fly.io’s Sprites disposable cloud containers now expose an MCP interface, making them a natural agent execution environment. [Fly.io]
- Agents that run while I sleep — Practitioner writeup on building persistent background agents with Claude Code; 429 HN points. [HN]
- Ask HN: How is AI-assisted coding going for you professionally? — High-signal community thread on real-world AI-coding outcomes; 434 points, 616 comments. [HN]
- Don’t post generated/AI-edited comments — HN is for conversation between humans — HN’s formal anti-AI-generated-content guideline update; 4,229 HN points — the week’s top item by score. [HN]
2026-W10
- A GitHub Issue Title Compromised 4k Developer Machines — “Clinejection” attack: a single maliciously crafted GitHub issue title triggered a popular AI coding tool to silently install a second agent on over 4,000 developer machines; 632 HN points. [HN]
- Codex Security: now in research preview — AI application-security agent that analyzes project context to detect, validate, and patch complex vulnerabilities with lower noise than traditional SAST. [OpenAI]
- Introducing GPT-5.4 — OpenAI’s 1M-token context frontier model with state-of-the-art coding, computer use, and tool search; 1,019 HN points. [OpenAI]
- Tell HN: I’m 60 years old. Claude Code has re-ignited a passion — Week’s top HN thread on agentic coding tools’ human impact; 1,086 points, 989 comments. [HN]
- If AI writes code, should the session be part of the commit? — Memento project proposes embedding full AI session transcripts in commits for auditability; 497 HN points. [HN]
- LLMs work best when the user defines acceptance criteria first — Practical framing: specify pass/fail criteria up front, treat AI output like a contractor deliverable; 461 HN points. [HN]
- Hardening Firefox with Anthropic’s Red Team — Anthropic red team partners with Mozilla to identify and harden Firefox security vulnerabilities; 629 HN points. [Anthropic / Mozilla]
- A standard protocol to handle and discard low-effort AI-generated pull requests — 406 HTTP-inspired rejection convention for AI-slop PRs; 305 HN points. [HN]
2026-W09
- What Claude Code chooses — Amplifying.ai research on which tools, patterns, and strategies Claude Code selects in practice; 611 HN points. [HN]
- Claude Code Remote Control — Docs for Claude Code’s remote control API, enabling external orchestration of the agent loop; 544 HN points. [Anthropic]
- MCP server that reduces Claude Code context consumption by 98% — Context-mode MCP server surgically limits context fed to Claude Code; 570 HN points. [HN]
- How we rebuilt Next.js with AI in one week — Cloudflare engineers use AI to rebuild a Next.js-compatible framework (Vinext) in one week; 540 HN points. [HN / Cloudflare]
- Why OpenAI no longer evaluates SWE-bench Verified — SWE-bench Verified declared contaminated; SWE-bench Pro recommended as replacement coding benchmark. [OpenAI]
- OpenAI Codex and Figma launch seamless code-to-design experience — Codex integration connecting implementation and Figma canvas for tighter design-to-code loops. [OpenAI]
- Best practices shipping multiple iOS apps with Claude Code — Community writeup of production patterns for mobile development with agentic coding; 434 upvotes. [r/ClaudeAI]
2026-W08
- Claude Sonnet 4.6 — Anthropic’s Sonnet release topping coding performance metrics; 1,346 HN points, 1,226 comments — week’s top HN post. [Anthropic]
- How I use Claude Code: Separation of planning and execution — Practical workflow post on keeping planning and implementation phases distinct; 976 HN points, 591 comments. [HN]
- Claws are now a new layer on top of LLM agents — Karpathy’s architectural framing for a tool-use controller layer above raw model calls; 412 HN points, 941 comments. [HN / Karpathy]
- I cut Claude Code’s token usage by 65% by building a local dependency graph and serving context via MCP — Community-built MCP context optimization; 183 upvotes. [r/ClaudeAI]
- Zed: Split Diffs are Here — Zed ships side-by-side diff view for AI-generated patches. [Zed]
- Monty: A minimal, secure Python interpreter in Rust for use by AI — Pydantic ships a sandboxed Python interpreter as a safe execution layer for LLM-generated code; 323 HN points. [Pydantic]
- Anthropic officially bans using subscription auth for third-party use — Policy update blocking OAuth credential reuse across third-party apps; 655 HN points. [Anthropic]
2026-W07
- Harness engineering: leveraging Codex in an agent-first world — Ryan Lopopolo (OpenAI) on how evaluation harness design shapes model coding performance more than the model itself. [OpenAI]
- Improving 15 LLMs at coding in one afternoon — only the harness changed — Changing evaluation harness design boosted 15 different LLMs without touching the models; 832 HN points, 295 comments. [HN]
- Introducing GPT-5.3-Codex-Spark — OpenAI’s first real-time coding model; 15x faster generation than its predecessor, 128k context window; research preview for Pro users; 890 HN points. [OpenAI]
- Custom Kernels for All from Codex and Claude — Hugging Face walkthrough on using Codex and Claude to write custom CUDA kernels. [Hugging Face]
- GLM-5 is on NVIDIA NIM — power Claude Code for free — Community project wiring GLM-5 on NIM as a drop-in Claude Code backend; 136 upvotes. [GitHub / r/LocalLLaMA]
- Warcraft III Peon voice notifications for Claude Code — Community notification tool; 1,006 HN points — a marker of Claude Code’s deep embedding in developer culture. [GitHub / HN]
- Claude Code is being dumbed down? — Community investigation into apparent Claude Code capability regression; 1,085 HN points, 701 comments. [HN]
2026-W06
- Orchestrate teams of Claude Code sessions — Official Anthropic docs for multi-session Claude Code orchestration; 396 HN points. [Anthropic]
- Claude Code is suddenly everywhere inside Microsoft — The Verge reports organic engineering adoption across Microsoft teams and Notepad integration; 409 HN points. [The Verge]
- Introducing GPT-5.3-Codex — OpenAI’s coding-focused model ships with a new macOS Codex app; 1,530 HN points, 605 comments. [OpenAI]
- Unlocking the Codex harness: how we built the App Server — OpenAI engineering deep-dive on the Codex app execution model. [OpenAI]
- Claude Code for Infrastructure (Fluid) — Startup building Claude Code as a layer for infrastructure provisioning; 276 HN points. [Fluid]
- Zed: Choose Your Edit Prediction Provider — Zed adds pluggable edit prediction backends, letting teams bring their own model. [Zed Blog]
- Zed: Community shipped 66 Git improvements in under 2 months — Open-source velocity benchmark for a production editor with heavy AI integration. [Zed Blog]
- Monty: A minimal, secure Python interpreter in Rust for use by AI — Pydantic ships a sandboxed Python interpreter designed as a safe execution layer for LLM-generated code; 323 HN points. [Pydantic]
2026-W05
- A few random notes from Claude coding quite a bit last few weeks — Karpathy’s real-world agentic-coding observations sparked the week’s biggest HN thread; 913 HN points, 848 comments. [HN / Karpathy]
- Porting 100k lines from TypeScript to Rust using Claude Code in a month — Detailed account of a full large-scale migration driven by Claude Code; 255 HN points. [HN]
- Claude Code daily benchmarks for degradation tracking — Community-built tracker measuring Claude Code output quality over time; 760 HN points, 354 comments. [HN]
- The ACP Registry is Live — Zed launches a public registry for Agent Cooperation Protocol (ACP) agents. [Zed Blog]
- How AI assistance impacts the formation of coding skills — Anthropic Fellows study finding AI assistance correlates with slower independent skill formation; 482 HN points. [Anthropic]
- There is an AI code review bubble — Greptile co-founder argues the AI code review tooling market is over-crowded and heading for consolidation; 351 HN points. [HN]
- Vibe coding is now just…coding — 560-upvote thread arguing AI-assisted coding has normalized enough to drop the “vibe” qualifier. [r/ChatGPTCoding]
2026-W04
- Claude Code’s new hidden feature: Swarms — Discovery post confirming Claude Code supports parallel sub-agent (swarm) orchestration; 521 HN points, 335 comments. [HN]
- Running Claude Code dangerously (safely) — Practical guide for running Claude Code with
--dangerously-skip-permissionsinside an isolated Docker container; 351 HN points. [HN] - I was banned from Claude for scaffolding a Claude.md file? — Developer account suspension story sparking 638-comment HN thread on automated policy enforcement and false-positive risk with agentic tooling; 749 HN points. [HN]
- Claude code creator shares update of v2.1.9 and hooks option — Boris Cherny posts update on Claude Code v2.1.9 including the new hooks system for lifecycle event handling; 110 upvotes. [r/ClaudeAI]
- Unrolling the Codex agent loop — OpenAI technical deep-dive into the Codex CLI agent loop: model orchestration, tools, prompts, and Responses API integration. [OpenAI]
- Show HN: Sweep — Open-weights 1.5B model for next-edit autocomplete — 1.5B open-weights model specialized for next-edit (diff-style) autocomplete; 534 HN points. [HN]
- On Programming with Agents — Zed’s philosophy piece on agents handling mechanical typing so developers can focus on higher-level reasoning. [Zed Blog]
2026-W03
- Cowork: Claude Code for the rest of your work — Anthropic’s research preview extending Claude Code-style agentic workflows to general office tasks; 1,298 HN points. [Anthropic]
- Claude Cowork exfiltrates files — PromptArmor documents a prompt-injection vector causing Claude Cowork to upload local files; 870 HN points — a landmark early disclosure for agentic AI security. [PromptArmor]
- First impressions of Claude Cowork — Simon Willison’s hands-on notes covering capabilities, limitations, and the file-exfiltration finding. [HN / Simon Willison]
- Show HN: OpenWork – An open-source alternative to Claude Cowork — Community responds to Cowork’s launch with an OSS alternative within 48 hours; 231 HN points. [GitHub]
- We put Claude Code in Rollercoaster Tycoon — Ramp Labs demo letting Claude Code play and modify RollerCoaster Tycoon — a stress test of agentic coding in a game environment. [HN / Ramp]
- My favorite command: /tutorial — Cursor’s
/tutorialcommand generates project-specific onboarding guides inline. [r/cursor]
2026-W02
- NVIDIA brings agents to life with DGX Spark and Reachy Mini — DGX Spark + Reachy Mini embodied-agent stack. [Hugging Face / NVIDIA]
2026-W01
- 2026 dev job market is straight-up cooked — 450-upvote r/cursor thread on the 2026 hiring-market reset. [r/cursor]
- 2 guys, 0 coding background, 10B tokens — 300k-line automotive platform with Cursor — Extended vibe-coding build retrospective; 128 upvotes. [r/cursor]
- Adaptive-P: a new sampler for creative text generation (llama.cpp PR) — Community sampler; 118 upvotes. [r/LocalLLaMA]
2025-W52
- What’s your best Claude Code non-coding use case? — Community catalog of non-coding Claude Code workflows; 104 upvotes. [r/ClaudeAI]
- My year with ChatGPT — 434-upvote retrospective; end-of-year community mood. [r/ChatGPTCoding]
2025-W51
- CUGA on Hugging Face: Democratizing Configurable AI Agents — IBM Research’s CUGA agent framework on the Hub. [Hugging Face]
- Tokenization in Transformers v5: Simpler, Clearer, and More Modular — Tokenizer refactor shipping with Transformers v5. [Hugging Face]
2025-W50
- New in llama.cpp: Model Management — First-class model downloads/caching/metadata in llama.cpp. [Hugging Face]
- Is this how Anthropic fixes major Claude outages? — 562-upvote r/ClaudeAI reliability discourse. [r/ClaudeAI]
- This is what happens when you vibe code so hard — Top r/ChatGPTCoding post (646 upvotes) on vibe-coding failure modes. [r/ChatGPTCoding]
2025-W49
- Claude CLI deleted my entire home directory — 1,538-upvote top thread on Claude Code’s rm-rf-~ incident and –dangerously-skip-permissions. [r/ClaudeAI]
- Transformers v5: Simple model definitions — Major Transformers release with cleaner model-definition API. [Hugging Face]
- Why I prefer Composer-1 as a senior software engineer — Senior-engineer take on Cursor’s Composer-1; 118 upvotes. [r/cursor]
- If You’re on 20x, put this in your CLAUDE.md — Community operator-playbook thread; 148 upvotes. [r/ClaudeAI]
2025-W48
- I compiled 2,000+ lines of Cursor tips and .cursorrules into one repo — Community-curated Cursor knowledge repo; 78 upvotes. [r/cursor]
- Continuous batching from first principles — HF explainer on inference-serving architecture. [Hugging Face]
2025-W47
- Google Antigravity — Agent-first IDE (VS Code fork); supports Gemini 3, Claude 4.6, and GPT-OSS-120B; free public preview. [Google Developers]
- GPT-5.1-Codex-Max — Long-horizon software-engineering model; first OpenAI model trained for Windows. [OpenAI]
- AnyLanguageModel: One API for Local and Remote LLMs on Apple Platforms — Unified local+remote LLM API for Apple. [Hugging Face]
2025-W46
- I tested GPT 5.1 Codex against Sonnet 4.5 — Early head-to-head; 93 upvotes. [r/ChatGPTCoding]
- Obsidian + Claude = Game Changer — Power-user workflow integration; 201 upvotes. [r/ClaudeAI]
2025-W45
- I built an entire fake company with Claude Code — 559-upvote agentic-construction case study. [r/ClaudeAI]
- Using Claude Code heavily for 6+ months: faster codegen hasn’t improved velocity — 400-upvote retrospective on why AI-accelerated codegen doesn’t mechanically translate to throughput. [r/ClaudeAI]
- Faster Prompt Processing in llama.cpp: Smart Proxy + Slots + Restore — Community prompt-processing optimization; 77 upvotes. [r/LocalLLaMA]
2025-W44
- The Claude Code hangover is real — 477-upvote thread on Claude Code quality drift. [r/ClaudeAI]
- Flutter 3.7 → 3.20 migration with Claude Code: saved $8k — Concrete migration success story; 204 upvotes. [r/ClaudeAI]
- Lazy Bird v1.0 — Automation harness that runs Claude Code on your projects while you’re at a day job; 110 upvotes. [GitHub]
2025-W43
- ChatGPT Atlas — OpenAI’s Chromium-based AI browser (macOS first) with sidebar ChatGPT and Agent Mode for multi-site action chains. [OpenAI]
- code-supernova-1-million — Stealth 1M-context coding model in Cursor and Cline. [Cursor / Cline]
- “You’re absolutely right!” is a bug, not a feature — Framing sycophantic agreement as model failure mode; 59 upvotes. [r/ClaudeAI]
- The prompt I run every time before git push (Codex or Claude Code) — Community pre-push prompt pattern; 64 upvotes. [r/ChatGPTCoding]
2025-W42
- Claude Skills — Composable, persistent AI capability framework for Claude; Willison calls it potentially bigger than MCP; 816 HN points. [Anthropic]
- Claude Skills: maybe bigger than MCP — Simon Willison’s analysis of why Skills changes the developer model for building on Claude; 738 HN points. [simonwillison.net]
- Beliefs true for software but false for AI — Framework for where standard engineering intuitions systematically break when applied to AI systems; 537 HN points. [boydkane.com]
- My trick for consistent LLM classification — Practical technique for reliable classification outputs from LLMs; 318 HN points. [Substack]
2025-W41
- Two things LLM coding agents are still bad at — Specific catalogued failure modes: context management and incremental refactoring at scale; 345 HN points. [kix.dev]
- Edge AI for Beginners — Microsoft’s curriculum for edge AI deployment on constrained hardware; 184 HN points. [GitHub]
- Recall: Redis-backed persistent context for Claude — Persistent memory layer for long-running Claude sessions; 171 HN points. [npm]
2025-W40
- Unix philosophy makes Claude Code amazing — Why composable shell-native design makes Claude Code fundamentally powerful; 414 HN points. [alephic.com]
- Comprehension debt: LLM-generated code time bomb — Teams shipping code they don’t understand accumulate cognitive debt that compounds; 532 HN points. [codemanship]
- Managing context on the Claude Developer Platform — Anthropic’s guidance on context window management for long-running agents; 220 HN points. [Anthropic]
- What ACTUALLY works after testing every AI coding tool for 6 months — Community synthesis of real-world AI coding tool effectiveness; 248 upvotes. [r/ChatGPTCoding]
2025-W39
- Getting AI to work in complex codebases — Stanford-derived advanced context engineering: explicit state management, tool orchestration, structured memory; 517 HN points. [GitHub]
- The AI coding trap — AI-generated code creates hidden comprehension debt that compounds over time; 685 HN points. [HN]
- Zed token-based LLM pricing — Zed shifts from flat-rate to consumption pricing for AI features; 182 HN points. [Zed]
- Pairing with Claude Code to rebuild a startup website — Practical Claude Code production case study; 178 HN points. [HN]
2025-W38
- GPT-5-Codex upgrades — IDE extensions, cloud integration, CLI enhancements; 396 HN points. [OpenAI]
- AI makes seniors stronger, not juniors — Why AI amplifies experienced developers’ advantage; 461 HN points. [HN]
- FAANG engineers using AI for production code — Real-world practices; 566 upvotes. [r/ChatGPTCoding]
2025-W37
- Geohot on AI coding — Geohot’s take on the state of AI coding; 410 HN points. [HN]
- Claude Code subagents for parallel development — Practical parallelization guide; 288 HN points. [HN]
- Claude’s memory architecture vs ChatGPT’s — Architecture comparison; 448 HN points. [HN]
2025-W36
- Claude Code in Zed via ACP — Zed integrates Claude Code using Agent Client Protocol; 683 HN points. [Zed]
- Claude Code modernizes 25-year-old kernel driver — Legacy system modernization; 929 HN points. [HN]
- Where’s the shovelware? — Why AI coding hasn’t produced predicted flood of apps; 770 HN points. [Substack]
- Staff engineer’s journey with Claude Code — “First attempt will be 95% garbage”; 550 HN points. [Sanity]
- AI-generated Metal kernels speed up PyTorch on Apple — AI-gen GPU kernels for Apple Silicon; 187 HN points. [HN]
2025-W35
- Claude Sonnet ships in Xcode 26 — Apple integrates Claude Sonnet natively; 489 HN points. [Apple]
- Claude for Chrome — Browser automation research preview for Max users; 799 HN points. [Anthropic]
- Grok Code Fast 1 — xAI enters the coding tools market; 512 HN points. [xAI]
- Martin Fowler on LLMs and software development — Authoritative essay; 420 HN points. [martinfowler.com]
- Nx supply chain attack exploits Claude Code CLI — First supply chain attack leveraging AI coding tools; 493 HN points. [Semgrep]
- 1/3 of senior devs say >50% code is AI-generated — Fastly survey; 215 HN points. [Fastly]
2025-W34
- What makes Claude Code so damn good — Technical deep-dive on Claude Code’s architecture; 469 HN points. [MinusX]
- Vibe coding creates a bus factor of zero — Essay on knowledge-transfer risks of AI-generated codebases; 223 HN points. [HN]
- Making games: 3 months without vs 3 days with LLMs — Concrete productivity comparison; 358 HN points. [HN]
- Tidewave Web: browser coding agent for Rails/Phoenix — In-browser coding agent for web frameworks; 299 HN points. [HN]
- Alibaba Qoder vs Cursor — Chinese AI coding tools emerging as competitors. [r/cursor]
2025-W33
- Claude Code is all you need — Essay arguing Claude Code is a sufficient development environment; 851 HN points. [HN]
- Why LLMs can’t really build software — Zed’s contrarian essay on AI coding limitations; 862 HN points. [Zed]
- Claude sycophancy: “You’re absolutely right!” — Viral issue on Claude’s over-agreement behavior; 773 HN points. [GitHub]
- Claudia — desktop companion for Claude Code — Desktop GUI wrapper; 501 HN points. [HN]
- Omnara — run Claude Code from anywhere — Remote execution tool; 310 HN points. [GitHub]
2025-W32
- Claude Code IDE integration for Emacs — Community Emacs integration; 782 HN points. [GitHub]
- Getting good results from Claude Code — Practical usage guide; 490 HN points. [HN]
- How I code with AI on a budget/free — Guide to free AI coding workflows; 697 HN points. [HN]
- AI 10x engineer imposter syndrome — Essay on AI coding’s psychological impact on developers; 949 HN points. [HN]
2025-W31
- Claude Code weekly rate limits — Anthropic announces new weekly rate limits for Claude Pro and Max; 609 HN points, 705 comments — triggers major community pricing debate. [HN]
- 6 Weeks of Claude Code — Puzzmo’s detailed retrospective on using Claude Code in production game development; 581 HN points. [Puzzmo]
- Crush: terminal AI coding agent — Charmbracelet launches Crush, a glamorous terminal-native AI coding agent; 367 HN points. [GitHub]
- Cerebras Code — Cerebras enters the AI coding space with its own product; 449 HN points. [Cerebras]
- Codestral 25.08 — Mistral’s enterprise coding stack claiming 50% reduction in dev/review/test time. [Mistral]
- Karpathy shifts to 80% agent-driven coding — Signals the tipping point for AI-assisted development workflows. [x.com]
2025-W30
- Claude Code specialized sub-agents — Claude Code gains specialized sub-agent capability for spawning purpose-built agent workflows inside a parent session. [Anthropic]
- How Anthropic teams use Claude Code — Anthropic publishes internal case studies showing how its own teams run Claude Code at scale. [Anthropic]
- Zed adds a toggle to disable all AI features — Zed’s single-toggle AI opt-out becomes a widely-shared marker for the coding-AI skeptic camp. [Zed]
- Replit AI agent wipes production database — First widely-reported catastrophic coding-agent failure in production. [BI]
- AI coding agents are removing programming language barriers — Rails at Scale essay on agents erasing language-choice constraints. [HN]
2025-W29
- Cognition acquires Windsurf — Cognition (Devin) acquires Windsurf’s brand, IP, enterprise business, and remaining staff 72 hours after Google’s $2.4B talent deal. Ends the fastest billion-dollar AI startup unwind on record. [Cognition]
- Anthropic silently tightens Claude Code limits — First real trust crisis for Claude Code subscribers. [TechCrunch]
- Claude for Financial Services — Anthropic’s regulated-vertical launch. [Anthropic]
- Conductor — Mac app that runs many Claude Codes at once — Multi-session Claude Code orchestrator. [HN]
- AI slows down OSS developers — Peter Naur can teach us why — Grounds the METR findings in “Programming as Theory Building.” [HN]
2025-W28
- METR study: AI tools slow down experienced OSS devs — Controlled study finding early-2025 coding tools made experienced OSS devs slower despite the devs’ beliefs. First widely-cited empirical counter to the “2x faster” claim. [METR]
- Google-Windsurf $2.4B licensing deal — Google hires Windsurf CEO, co-founder, and ~40 senior R&D staff in a licensing deal. [TechCrunch]
- Running Claude Code inside Docker + VS Code — Guide for sandboxing Claude Code; becomes the reference pattern. [HN]
- MCP-B: Protocol for AI browser automation — Proposed MCP extension for browser automation. [HN]
- Morph (YC S23): apply AI code edits at 4,500 tokens/sec — Fast code-edit-application model targeting agent latency. [HN]
- Vibe Kanban — Kanban UI for orchestrating multiple coding agents. [HN]
- BentoML LLM Inference Handbook — Complete handbook on serving LLMs at scale. [HN]
2025-W27
- Context engineering > prompting — Philipp Schmid’s essay reframes agent design as long-horizon state management. 915 HN points — the week’s top post. [Phil Schmid]
- Claude Code ships hooks + goes GA — Event-driven hooks enable custom guardrails, observability, and policy integrations. [Anthropic]
- Spegel: LLM terminal browser that rewrites webpages — New primitive for LLM-augmented reading. [HN]
- I’m dialing back my LLM usage — Senior dev interview on scaling down LLM reliance. [Zed]
- What to build instead of AI agents — Counter-argument for workflows over agents. [HN]
- Hamel Husain: About AI evals — Canonical FAQ on AI evals. [HN]
- Building a Personal AI Factory — Practitioner pipeline for personal AI-augmented coding. [HN]
2025-W26
- Gemini CLI — Google’s open-source terminal coding agent, the direct answer to Claude Code. Week’s top post at 1,428 HN points. [Google]
- Claude Code for VSCode (GA) — Claude Code GA VS Code extension. [Anthropic]
- Claude Artifacts become hostable apps — No-deployment AI-powered apps built inside Claude. [Anthropic]
- QEMU bans AI code generators — First major OSS project to formally forbid AI-generated contributions. [HN]
- Learnings from building AI agents (Cubic) — Practical production-agent lessons. [HN]
- Magnitude: open-source AI browser automation framework — OSS agentic browser automation. [HN]
2025-W25
- Karpathy: Software in the era of AI — Andrej Karpathy’s YC Startup School keynote introducing “Software 3.0”; becomes the dominant vocabulary for coding-AI discourse through Q3. [YouTube]
- Anthropic: Building Effective AI Agents — Canonical engineering guide for production agents. [Anthropic]
- Remote MCP Support in Claude Code — First frontier-vendor coding agent with remote MCP. [Anthropic]
- Phoenix.new — remote AI runtime for Phoenix — Fly.io’s coding-agent-optimized Phoenix runtime. [Fly.io]
- Snorting the AGI with Claude Code — Practitioner chronicle of heavy Claude Code Max usage. [HN]
- Nxtscape: open-source agentic browser — OSS agentic browser alternative. [HN]
2025-W24
- Claude Code with 20M free tokens loophole — Community-discovered Claude Code Max usage loophole. [r/ClaudeAI]
- Tritium: the Legal IDE in Rust — AI-first IDE aimed at lawyers. [HN]
- A “Course” as an MCP Server — Mastra ships a course consumed over an MCP server. [HN]
- Chonkie: open-source advanced chunking library — Chunking library for long-context RAG pipelines. [HN]
- Fine-tuning LLMs is a waste of time — Contrarian argument for context engineering over fine-tuning. [HN]
2025-W23
- Cloudflare builds OAuth with Claude and publishes all the prompts — Best-documented public case of shipping production infrastructure via a frontier model. [Cloudflare]
- I read all of Cloudflare’s Claude-generated commits — Companion commit-by-commit analysis. [HN]
- Mistral Code — Mistral’s agentic coding assistant built on Devstral + Codestral. [Mistral]
- Tracking Copilot vs Codex vs Cursor vs Devin PR performance — Public leaderboard comparing coding agents on real PRs. [HN]
- Field notes from shipping real code with Claude — Practitioner write-up on production Claude coding. [HN]
- Reverse engineering Cursor’s LLM client — Deep dive on Cursor’s internals and model routing. [HN]
- “My AI skeptic friends are all nuts” (Thomas Ptacek / fly.io) — The week’s defining pro-coding-AI essay; 2,356 HN points. [Fly.io]
2025-W22
- antirez: human coders are still better than LLMs — Redis creator’s long-form pushback against post-Claude 4 coding-AI hype; 655 HN points. [HN]
- AI: Accelerated Incompetence — Argument that AI coding tools let junior engineers entrench bad patterns. [HN]
- Simon Willison: llm CLI gains tool use — Willison’s
llmCLI adds Python-plugin tool use. [HN] - Simon Willison: highlights from the Claude 4 system prompt — Annotated walkthrough of the leaked Claude 4 system prompt. [HN]
- Onlook: open-source visual-first Cursor for designers — Design-oriented alternative to Cursor. [HN]
- AutoThink: adaptive-reasoning harness for local models — Adaptive test-time reasoning harness. [HN]
2025-W21
- Claude Code SDK — Programmatic API for scripting Claude Code as an embeddable coding agent; 454 HN points. [Anthropic]
- Devstral — Mistral’s 24B open-weight coding model claiming SOTA on SWE-Bench Verified among open-weight models. [Mistral]
- Google AI Studio native code generation and agentic tools upgrade — Google AI Studio gains code generation and agent tooling. [Google]
- Trading with Claude and writing your own MCP server — Practitioner build-log on custom MCP servers for finance. [HN]
- Peer programming with LLMs, for senior+ engineers — Practitioner patterns for pairing with LLMs at senior level. [HN]
2025-W20
- The unreasonable effectiveness of an LLM agent loop with tool use — David Crawshaw’s essay arguing minimal LLM agent loops plus tool use beat elaborate agent frameworks; 447 HN points. Canonical coding-agent reference. [HN]
- Getting AI to write good SQL — Google Cloud’s comprehensive text-to-SQL quality techniques; 501 HN points. [Google]
- LLM function calls don’t scale; code orchestration is simpler, more effective — Argues code orchestration beats function-calling for large-data workflows. [HN]
- KVSplit: 2-3x longer contexts on Apple Silicon — KV-cache splitting for M-series chips. [HN]
- Muscle-Mem: behavior cache for AI agents — Caches agent trajectories to avoid repeat inference. [HN]
- Airweave: let agents search any app — Universal search API for agent-accessible SaaS. [HN]
- After months of coding with LLMs, I’m going back to my brain — Senior dev’s LLM-coding burnout post-mortem. [HN]
2025-W19
- Zed: high-performance AI code editor — Zed repositions as a performance-first AI editor; direct shot at Cursor and VS Code on raw responsiveness. [HN]
- Claude’s 24k-token system prompt leaked — Full reverse-engineered dump of Claude’s system prompt including tool definitions. [HN]
- Claude Code flat pricing via Max plan — Claude Code folded into $100/$200 Max subscriptions. First credible alternative to Cursor’s per-seat pricing. [Anthropic]
- Notes on rolling out Cursor and Claude Code — Practitioner write-up of enterprise rollout. [HN]
- Real-time AI voice chat at ~500ms latency — Open-source real-time voice agent. [HN]
- Nao Labs: Cursor for data — AI IDE targeted at data engineers. [HN]
- Mike Krieger: 70%+ of Anthropic PRs are AI-generated — First-party evidence from a frontier lab that AI coding works at scale. [r/ClaudeAI]
2025-W18
- Chain of Recursive Thoughts — Novel prompting technique using internal debate for improved reasoning; 539 HN points. [HN]
- Tiny-LLM: serving LLMs on Apple Silicon — Systems engineering course for M-series chips; 297 HN points. [HN]
- Anemll: LLMs on Apple Neural Engine — Framework for running LLMs on Apple’s dedicated ML hardware; 286 HN points. [HN]
- RustAssistant: LLMs fixing Rust compilation errors — Microsoft Research on LLM-powered Rust compiler error resolution; 142 HN points. [HN]
- AI code review: author as reviewer conflict — Analysis of conflict-of-interest when AI reviews code it helped write; 136 HN points. [HN]
- Vibe coding era: billions of lines with millions of bugs — Community concern about code quality; 100 upvotes. [r/cursor]
2025-W17
- AI Horseless Carriages — Influential essay arguing current AI products replicate old workflows instead of building AI-native paradigms; 864 HN points. Became a meme for critiquing surface-level AI integrations. [HN]
- The hidden cost of AI coding — Analysis of tech debt from AI-generated code; 341 HN points, 462 comments. [HN]
- Avoiding skill atrophy in the AI age — Preserving developer skills while using AI; 373 HN points. [HN]
- LLM tools as mech suits, not replacements — “Amplification not replacement” framing; 345 HN points. [HN]
- Teaching LLMs solid modeling — Getting LLMs to generate valid 3D CAD; 319 HN points. [HN]
- AI assisted search-based research works now — Simon Willison declares AI search has crossed usefulness threshold; 283 HN points. [HN]
2025-W16
- Claude Code: Best practices for agentic coding — Anthropic’s definitive guide to CLAUDE.md, sub-agents, and test-driven agent workflows; 614 HN points. [Anthropic/HN]
- SQLite + cron AI assistant — Elegant minimal architecture for a personal AI assistant; 800 HN points. [HN]
- 12-factor Agents — Reliability patterns for LLM apps; 475 HN points. [HN]
- JetBrains IDEs Go AI — Native AI agent capabilities in all JetBrains IDEs; 174 HN points. [HN]
- AgentAPI for Claude Code, Goose, Aider — Unified HTTP API for multiple AI coding agents; 163 HN points. [HN]
- Cursor AI support hallucinates lockout policy — AI customer support fabricated a non-existent policy causing cancellations; 1,511 HN points. Cautionary tale for AI in customer-facing roles. [HN]
- AI that turns codebases into tutorials — AI-powered codebase documentation; 923 HN points. [HN]
2025-W15
- Browser MCP — MCP server for browser automation via Cursor/Claude/VS Code; 616 HN points. [HN]
- AI coding PB&J problem — AI’s inability to handle implicit context; 122 HN points. [HN]
- Chonky: neural text semantic chunking — Neural chunking outperforming rule-based for RAG; 169 HN points. [HN]
- Karpathy’s AI coding rhythm — “Keep a tight leash on this over-eager junior intern savant” framing for AI-assisted coding workflows. [x.com]
2025-W14
- Senior Developer Skills in the AI Age — Essay on what makes experienced developers effective with AI agents; 421 HN points, 318 comments. [HN]
- Interviewing a software engineer who prepared with AI — Kapwing’s account of AI-coached candidates passing screening but failing to explain answers; 415 HN points, 708 comments — signals early concern about AI masking skill gaps in hiring. [HN]
- Don’t let an LLM make decisions or execute business logic — Production-focused essay on constraining LLM scope; 325 HN points, 169 comments. [HN]
- Pydantic Evals: evaluation framework for LLM projects — Pydantic AI team releases dedicated eval tooling; lowers barrier for production LLM app evaluation. [x.com/@simonw]
- Show HN: Qwen-2.5-32B is now the best open source OCR model — OmniAI benchmark showing Qwen-2.5-32B topping open-source OCR; 211 HN points. [HN]
- LLM Workflows then Agents: Airflow AI SDK — Astronomer’s SDK for building LLM pipelines on Apache Airflow; applies battle-tested DAG orchestration to agentic workloads. [HN]
2025-W13
- The role of developer skills in agentic coding — Martin Fowler’s series argues developer fundamentals remain essential in agentic coding; 336 HN points, 195 comments — high-signal community discussion on what skills actually matter. [HN]
- Launch HN: Continue (YC S23) – Create custom AI code assistants — Continue launches a hub for shareable custom AI coding assistants built on their open-source VS Code/JetBrains extension. [HN]
- Kilo Code: Speedrunning open source coding AI — New open-source Cursor/Copilot competitor aiming for fast iteration on coding AI features. [HN]
- Show HN: Cursor IDE now remembers your coding prefs using MCP — MCP-based extension to persist user coding preferences across Cursor sessions; signals MCP expanding beyond data sources into UX customization. [HN]
- Learn to code, ignore AI, then use AI to code even better — Essay arguing foundational coding skills amplify rather than compete with AI tools; 158 HN points with substantive debate. [HN]
2025-W12
- Claude web search — Real-time web search with auto-triggering and direct citations for Claude 3.7 Sonnet. [Anthropic]
- AI Blindspots — Catalog of systematic LLM coding weaknesses: pattern matching failures, hallucinated APIs, context edge cases. [ezyang]
- How Cursor (AI IDE) Works — Technical deep-dive into Cursor architecture: prompt construction, indexing, tab completion, agent mode. [Blog]
- Fine-tune Gemma 3 with Unsloth — Efficient Gemma 3 fine-tuning guide. [Unsloth]
- Simon Willison: Vibe coding vs AI-assisted programming — Distinguishing throwaway vibe coding from responsible AI-assisted development. [Blog]
2025-W11
- Anthropic token-saving API updates — Token-efficient tool use (up to 70% reduction), cache-aware rate limits for Claude 3.7 Sonnet. [Anthropic]
- codemcp — MCP server giving Claude Desktop coding agent capabilities via Pro subscription. [HN]
- Aider v0.77.0 — Now supports 130 new programming languages. [Reddit]
- Mistral OCR — 2000 pages/minute OCR API with LaTeX/table/handwriting support. [Mistral]
- Xata Agent — AI agent specialized in PostgreSQL queries and management. [HN]
2025-W10
- Microsoft 365 Copilot adds AI Sales Agents — Microsoft extends Copilot into specialized sales workflows, moving beyond developer tooling toward autonomous enterprise productivity agents. [Microsoft]
- LLM Inference on Edge via React Native — Hugging Face guide to running local LLMs in mobile apps via React Native; signals maturing mobile inference toolchain for developers. [Hugging Face]
Baseline (through 2025-Q1)
AI-Native IDEs and Editors
- Cursor — AI-first VS Code fork by Anysphere. Became the breakout IDE of 2024 with inline editing, multi-file diffs, and codebase-aware chat. Raised $100M+ and grew rapidly. Popularized “Cmd+K” inline editing UX pattern.
- Source: https://www.cursor.com/
- Windsurf (Codeium) — Codeium rebranded its IDE product to Windsurf (Nov 2024). Focused on “flows” — agentic multi-file editing with contextual awareness. Free tier included.
- Source: https://codeium.com/windsurf
- Zed — High-performance editor (Rust-based) added AI assistant features with support for multiple LLM providers. Open-source.
- Source: https://zed.dev/
Coding Assistants and Agents
- Claude Code (Feb 2025) — Anthropic’s agentic CLI coding tool. Runs in the terminal, can read/write files, run commands, search codebases, and manage git workflows. Designed for long-running autonomous coding tasks. Uses Claude 3.5 Sonnet as default model.
- GitHub Copilot — Still the largest installed base. Added Copilot Chat, multi-file editing (Copilot Edits), and agent mode. Switched from Codex to GPT-4o-based models. Copilot Workspace (preview) offered issue-to-PR automation.
- Devin (Cognition Labs, Mar 2024) — Marketed as “first AI software engineer.” Browser-based autonomous coding agent with shell, editor, and browser access. Generated enormous hype. Raised $175M at $2B valuation. Real-world performance debated.
- Aider — Open-source terminal-based AI pair programming tool. Supports multiple LLMs. Strong SWE-bench performance. Pioneered “architect” mode using two models (one to plan, one to edit).
- Source: https://aider.chat/
- Amazon Q Developer — AWS’s coding assistant (rebranded from CodeWhisperer). Focused on AWS service integration and enterprise compliance.
- Google Jules (Dec 2024) — Google’s AI coding agent, async task execution on GitHub issues/PRs. Powered by Gemini 2.0.
Code-Specialized Models
- Qwen 2.5 Coder 32B (Nov 2024) — Open-weights coding model matching GPT-4o on code benchmarks. Became popular for local coding setups.
- DeepSeek-Coder-V2 — Open MoE model strong on code. Precursor to DeepSeek-V3’s coding capabilities.
- Source: https://arxiv.org/abs/2406.11931
- Codestral (Mistral, May 2024) — 22B code-specific model from Mistral, designed for code completion and generation.
Developer Productivity Research
- Google internal study (2024) — Reported developers using AI completed tasks 20-30% faster. But noted quality concerns with AI-generated code in code review.
- Stack Overflow Developer Survey 2024 — 76% of developers using or planning to use AI coding tools. Trust in accuracy was mixed, with only 43% trusting AI output “a great deal” or “to a moderate extent.”
- SWE-bench — Standard benchmark for coding agents. Verified subset became the gold standard. Top agents (using Claude 3.5 Sonnet) exceeded 50% resolution rate.
- Source: https://www.swebench.com/