Claude Opus 4.7 Dropped — And a Local Model Drew the Better Pelican

Claude Opus 4.7 Dropped — And a Local Model Drew the Better Pelican

Claude Opus 4.7 is here with upgraded vision, memory, and instruction-following — but Simon Willison's pelican benchmark just handed the win to a local Alibaba model running on a laptop. We dig into what that actually means, plus Anthropic's new identity verification layer, Amazon's MCP bet, and whether "personal software" is about to change who gets to be a developer. Your commute just got more interesting.

Det här avsnittet är hämtat från ett öppet RSS-flöde och publiceras inte av Podme. Det kan innehålla reklam.

Avsnitt(25)

Passing Tests Isn't Enough for Your Next Coding Agent

Passing Tests Isn't Enough for Your Next Coding Agent

Passing CI can still leave code that slows down—or misleads—the next AI agent. Fictional AI hosts Alex and Sam use this week’s debate about Go and agent-friendly engineering to build a practical machi...

14 Aug 18min

Your OpenClaw Updates Need a Canary, Not Courage

Your OpenClaw Updates Need a Canary, Not Courage

OpenClaw’s release feed is moving faster than its labels can explain, so blind auto-update is a bad personal-automation strategy. Cleo and Dev build a Release Sentinel canary, keep telemetry local, an...

2 Aug 20min

Claude Code Changed Engines—Your Evals Just Broke

Claude Code Changed Engines—Your Evals Just Broke

Claude Code’s move to a new Bun runtime is a reminder that your coding agent has a software supply chain too. Alex and Sam unpack runtime drift, model routers, reverse-engineering with agents, and a f...

24 Juli 19min

Better Agent Tools Made Code Review Worse

Better Agent Tools Made Code Review Worse

GitHub gave its code-review agent better tools and watched cost rise while useful findings fell. Alex and Sam unpack why task-shaped instructions beat bigger toolboxes, how invisible environment detai...

14 Juli 18min

Your AI Coding Benchmarks Are Lying To You

Your AI Coding Benchmarks Are Lying To You

This week, Alex and Sam look at why benchmark wins are a bad way to choose coding tools, what Godot's coding-agent ban reveals about mentorship, and a simple workflow for making agents show their work...

3 Juli 18min

The Tiny Local Model That Changes Your Agent Budget

The Tiny Local Model That Changes Your Agent Budget

Small, local models are suddenly good enough for real agent chores, but the win is not replacing your smartest model. Cleo and Dev unpack lightweight extraction models, model-routing memory, browser-s...

26 Juni 18min

Your Coding Agent Needs a Bouncer Now

Your Coding Agent Needs a Bouncer Now

AI coding agents are getting longer runs, more context, and more ways to touch production workflows, but this week made the real bottleneck obvious: authorization. Alex and Sam unpack MCP's missing en...

19 Juni 19min

Verification Is Now Your Coding Agent Bottleneck

Verification Is Now Your Coding Agent Bottleneck

Coding agents are getting better at long runs, but this week's news points at the real limit: proof. Alex and Sam unpack agent loops, Stack Overflow for Agents, Copilot CLI delegation, local-model cod...

17 Juni 11min

Populärt inom Teknik

uppgang-och-fall
market-makers
skogsforum-podcast
rss-uppgang-och-fall
rss-laddstationen-med-elbilen-i-sverige
rss-elektrikerpodden
rss-en-ai-till-kaffet
natets-morka-sida
 och-bilen-gar-bra
bli-saker-podden
rss-veckans-ai
under-femton
hej-bruksbil
elbilsveckan
developers-mer-an-bara-kod
bosse-bildoktorn-och-hasse-p
rss-fabriken-2
garagehang
klocksnack-tillsammans-med-nymans-ur-1851
bilar-med-sladd