Migrating to Claude Opus 5.5: a cheat sheet for the API and Claude Code
Effort defaults, the four breaking changes, silent progress updates, and the prompt and CLAUDE.md lines worth changing when you move from Opus 5 to 5.5.
Everything published here, newest first. Use ⌘Ctrl K to search.
Effort defaults, the four breaking changes, silent progress updates, and the prompt and CLAUDE.md lines worth changing when you move from Opus 5 to 5.5.
My open-source Agent Skills bundle gives coding agents a shared story bible format — and a deterministic CLI that catches dead characters walking, payoffs before setups, and unfired Chekhov guns.
Meta's Muse, xAI's Grok Bot, and the open-source route (OpenClaw and Hermes Agent) are all fighting to be your always-on AI assistant. They disagree on who does the setup, whose model you get, and where your data lives.
GPT-5.6 Sol, Kimi K3, and Claude Opus 5 trade coding wins. An early comparison of capability, cost, availability, and agent workflows.
Skills like Caveman and tools like RTK save tokens, but their benchmarks rarely measure whether the agent still completes the work correctly.
Google DeepMind shipped an open-weight model that generates text by denoising a block of tokens instead of predicting one at a time. Here's how it works and where it fits.
AISI's evaluation puts GPT-5.5 beside Claude Mythos on advanced cyber tasks. The response is faster patching, better logs, and stricter access.
A first look at Axocoatl, a new Rust runtime where persistent AI agents coordinate through pheromone-style signals instead of a central orchestrator.
The best AI coding results come from telling the model what you're trying to build, then steering the conversation until the shape is right.
Zero is a pre-1 systems language experiment where the compiler, standard library, and diagnostics are designed for AI agents first.
A deep look at the Claude Code plugin that turns codebases into interactive knowledge graphs — how the multi-agent pipeline works, what the dashboard looks like, and where the rough edges are.
A look at tw93's Mole — a CLI tool with over 40k GitHub stars that consolidates CleanMyMac, DaisyDisk, iStat Menus, and AppCleaner into a single free utility.
Vercel Labs reimplemented bash and 70+ Unix commands in pure TypeScript, giving AI agents a sandboxed shell that runs anywhere JavaScript does.
Three Chinese MoE models claim frontier-class coding at a fraction of Opus pricing. Here's how they actually perform in Claude Code, Cursor, and real developer workflows.
The difference between managing and micro-managing AI agents is the same gap that separates good tech leads from bad ones. Share the why, not just the what.
Introducing browse, a CLI tool that wraps Playwright behind a persistent Unix socket daemon for sub-30ms browser automation commands.
NVIDIA's Nemotron 3 Super packs 120B parameters into 12B active, combining Mamba-2, Transformers, and a novel LatentMoE — all open-weight and purpose-built for multi-agent systems.
cmux is a native macOS terminal built on Ghostty's rendering engine, designed for running multiple AI coding agents in parallel with rich notifications and a scriptable API.
notebooklm-py gives you programmatic access to Google NotebookLM. Even better, it works as an AI agent skill — so you can generate podcasts, quizzes, and slide decks just by asking Claude or ChatGPT.
Claude Code can poll deployments, babysit PRs, and remind you to push — all with a single /loop command or a plain-English sentence.
OpenAI's Responses API now supports WebSocket transport. For tool-call-heavy agent loops, it cuts end-to-end latency by up to 40%.
Kimi K2.5 delivers 95% of Opus 4.6's coding capability at 10-25× lower cost. But the benchmarks don't tell the whole story.
A practical guide to when Skills, CLI tools, and Model Context Protocol each shine — and why the best agent setups use all three.
Block cut 4,000 employees while posting record profit. The development pipeline has permanently shifted — and most companies haven't caught up.
How I built openclaw-watchdog — a lightweight monitoring tool that uses Claude Code to automatically diagnose and fix OpenClaw gateway failures.
AI coding agents forget everything between sessions. git-semantic-bun lets them search commit history by meaning — locally, offline, and in a format machines can consume.
A practical, reproducible guide to running OpenClaw on Proxmox with Ubuntu, including VM sizing, gateway service setup, security hardening, and troubleshooting.
Two new Anthropic studies reveal that users voluntarily cede judgment to AI — and feel good about it until things go wrong.
Google's latest frontier model more than doubles its predecessor's reasoning score in three months, leads 13 of 16 benchmarks, and ships at the same price. The adaptive compute architecture is the interesting part.
A detailed comparison of Claude Opus 4.6 and GPT-5.3-Codex for Laravel/Vue/Inertia stacks — benchmarks, costs, ecosystem depth, and practical recommendations.
Anthropic's second major launch in two weeks puts near-flagship capability at $3/$15 per million tokens. The mid-tier label is starting to feel like a misnomer.
A look at the most popular Claude Code plugin — how it turns AI coding agents from capable-but-undisciplined into structured development partners through composable skills.
A seven-step process for turning requirements into shipped features using Claude Code — from scoping with parallel sub-agents to test-driven execution.
Claude Code now has built-in auto memory, but a 28k-star plugin called claude-mem offers a far more sophisticated approach. Here's how they compare and when each makes sense.
A 230B MoE model with 10B active parameters hits 80.2% on SWE-Bench Verified at 1/20th the cost of Opus. Here's what's real and what's hype.
Zhipu AI's 744B mixture-of-experts model ships under MIT license with frontier-class benchmarks and aggressive pricing. Here's what actually matters.