Beyond vibe checks: A PM’s complete guide to evals
Explains why evaluation skill is becoming essential for building reliable AI products.
Pieces about AI, gathered from every area. Each one also appears on its own subject page.
Explains why evaluation skill is becoming essential for building reliable AI products.
Explains Netflix's shift from hand-crafted feature models to an LLM-native recommendation architecture.
Reports that an AI model found a counterexample disproving a 1946 Erdős conjecture.
Scrutinizes the widely cited claim that dynamic languages are more token-efficient for coding agents.
Describes a research-plan-implement workflow that withholds coding until a written plan is approved.
Asks whether AI 'reasoning' models reach correct answers through fundamentally flawed processes.
Argues that LLM capabilities stem mainly from imitative learning rather than reinforcement learning.
Explains how to build a self-improving PM workflow with Claude Code's skills and CLAUDE.md router pattern.
Explains the design rationale behind MCP 2.0's move to a stateless protocol.
Examines how AI coding agents strain code review practices built for human-scale output.
Walks through Claude Code's anatomy: the .claude directory, CLAUDE.md conventions, skills, and subagents.
Reports measured results from an AI code reviewer that flagged violations and blocked merges over four months.
Maps the full threat model for LLM security, using a 2025 Microsoft 365 Copilot data-exfiltration case.
Outlines four research-based ways analytical work can direct AI to generate surprising insights.
Compiles ten concrete results AI models have produced in mathematics and theoretical computer science.
Argues for making agent-driven engineering work falsifiable and resumable, via a compiler-age analogy.
Explains why teams that treat AI as shared infrastructure see bigger productivity gains than solo hackers.
Argues that human judgment doesn't disappear from automated software production, it moves elsewhere.
Lays out the full timeline of OpenAI's accidental agent-driven attack on Hugging Face.
Breaks down why serving LLM memory gets costly at scale and how engineering teams can reduce it.
Offers a skeptical, evidence-based account of using AI agents for coding tasks over several months.
Argues that new tools, not just ideas, have driven history's biggest leaps in scientific understanding.
Applies a who/what/where/when/why/how framework to critical thinking in the age of AI.
Works through benchmarking and evaluation puzzles, including SWE-Bench napkin math and DeepSWE.
Explains in depth how the Model Context Protocol works under the hood.
Weighs unresolved questions about the risks and benefits of releasing AI model weights openly.
Argues that AI-generated writing with no communicative intent produces prose nobody truly wants to read.
Argues that software engineering fundamentals matter more, not less, in the age of AI coding tools.
Argues for guarding independent judgment against the pull to offload thinking onto AI tools.
Walks through designing a vertical small LLM system to classify and route support tickets end to end.
Introduces the fundamentals of agentic, AI-assisted software development practices.
Explains how large models transfer knowledge to smaller ones through the mechanism of distillation.
Recounts firsthand how building an AI executive assistant with Claude Code reshaped a CTO's daily work.
Reports on an AI agent that attacked unrelated companies during a government cyber evaluation.
Surveys recurring architectural patterns and failure modes in emerging multi-agent AI systems.
Argues that AI will push formal verification from a fringe pursuit into mainstream engineering.
Argues the file system already behaves like a graph database for organizing personal knowledge.
Argues DAU, WAU and MAU are becoming the core metric for AI-native B2B products, citing Harvey.
Describes a personal workflow for planning, coding and structuring AI-assisted software development.
Explains why MCP is moving to a stateless connection model and what that changes in protocol design.