Kennis die lekker wegluistert
Onze podcasts - met liefde gemaakt door team Sourcelabs - en met (meer dan) een vleugje AI. Lekker voor onderweg.
Een dagelijkse AI-gegenereerde podcast over agentic AI, developer tooling en tech trends — volledig autonoom geproduceerd. Beschikbaar als RSS feed.
The Daily Agentic AI Podcast - 2026-10-09
The Daily Agentic AI Podcast - 2026-10-08
Anthropic released Claude Haiku 5.5, a cheaper small model with a million-token context window, adjustable effort settings, and adaptive thinking, though tokenizer changes and price jumps above 100k tokens complicate the headline savings. The episode also covered Anthropic's API credits, halved Sonnet 5.5 cache pricing, and SDK computer/browser use toolsets, alongside OpenAI's GPT-6 Intelligent UI rollout, Liquid AI's open-weight decision models, Stacklok's Mecatl agent harness, and Liquid Inference's per-request inference auction. It discussed Meta and Microsoft reducing internal Claude usage, Microsoft's hybrid Windows AI push, and research including LEGO code primitives, TestGRAD, concurrency tradeoffs in sub-agents, benchmark integrity audits, SpecGuard, NVIDIA's PivotOPD recovery distillation, AgentTracer, agentic constraint programming, market-driven alignment, and predicting post-training performance from base models.
The Daily Agentic AI Podcast - 2026-10-07
OpenAI released 722 mathematical manuscripts from an unreleased internal frontier model, prompting both major praise and skepticism about disproofs, scooping, and results surviving scrutiny. Google's FlowAgent repairs failing tests in pre-submit workflows, while a security benchmark found auto-approve pushed attack success in coding harnesses from under a third to nearly everything, and ParanoiaEval showed stronger capability does not predict better risk judgment. Other topics included Sign in with ChatGPT OAuth, Claude in Google Docs and Cloud sessions, Google Cloud API Gateway as a remote MCP server, harness engineering behavioral evaluations and Dev-Primitives, Agent Plugins 1.0.0, the EPIC framework, a healthcare software case study, GitHub's Git infrastructure rebuild, the Recursive Game Creator, CheckerBench, CLEAR, memory failures from re-reading old sources, PBT-Bench, corrupted tool feedback reliability, ADK live evaluation, Laya, Underdog funding, the functionality-security gap, brain-signal-guided fine-tuning, a Slack bank balance leak, autofinetune, zero-trust runtime governance, and failure-conditioned RL.
The Daily Agentic AI Podcast - 2026-10-06
Reflection AI announced Beam, its first open-weight 501B-parameter MoE model with 23B active parameters, Apache 2.0 weights due later in October, claiming three-to-four times less inference compute than GLM 5.2 at comparable reasoning scores, though Artificial Analysis found it competitive rather than dominant on coding and agentic benchmarks. Other model news included Mistral Large 4's research preview (1T total, 49B active, 500K context) and Reka's Rho-1 19B omni-reasoning research preview that outputs text, images, video, and robot actions from a shared latent state. On the tooling side, the episode covered Codemode (Pi's code-as-tool-call abstraction replacing MCP tool schemas with sandboxed JavaScript), Together Link (a CLI routing coding agents to open models), Antigravity SDK local model support via Gemma 4 26B and LiteRT, and the MCP stateless specification update removing transport-level sessions. Research segments covered a study finding harnesses mainly buy token cost savings rather than task competence, UndoBench-style findings that agent recovery rates lag completion rates, adaptive code revision attacks (AFCRA) defeating AI PR reviewers, GitHub ReviewBench, ADK for Kotlin 1.0, RepoLaunch's 78% build success rate, asymmetric repository lineage conflicts between concurrent agents, and a batch of safety papers including Wikimedia's confirmed OpenAI rogue agent activity, StateWise for stale operational state, IEC for intent-execution mismatches, Agent MechSuits subspace steering, simulate-vs-develop software comparison, and practitioner permission decision studies.
The Daily Agentic AI Podcast - 2026-10-05
Airbnb's CTO reported that roughly 60% of its code is now AI-authored, driving large gains in feature shipments, PR throughput, and AI-resolved support tickets, supported by an internal context graph and multi-model strategy. Aleph Alpha released Kolibri, an open-weight bilingual Apache-licensed mixture-of-experts model with 78B total and ~3.5B active parameters, a million-token context, and single-GPU serving. The episode also covered IBM's self-hosted air-gapped Bob platform, an Opus 5.5 Claude Code playbook, Google's argument for Go in AI-assisted engineering, generative test-driven development, a study showing multi-agent systems using over six times the energy of single-query baselines, decision pins for underspecified choices, a verify-before-you-fix cross-language vulnerability framework, Managed Deep Agents 0.8 user memory, a production incident failure taxonomy, Cantina's open security research model, Yandex's single generative recommender, and session-aware load balancing for long-lived agent sessions.
The Daily Agentic AI Podcast - 2026-10-04 - Special: The Bottleneck Was Never the Code: an extra deep dive on why AI makes coding faster but not shipping faster, with models converging, gated frontier releases, and the organisation as the real bottleneck
Special episode: The Bottleneck Was Never the Code: an extra deep dive on why AI makes coding faster but not shipping faster, with models converging, gated frontier releases, and the organisation as the real bottleneck. AI coding tools roughly triple commit volume but increase releases by only about 30%, because review, testing, and deployment remain human-speed bottlenecks that now cap the pipeline. The episode also examines converging frontier model capabilities, restricted access to the most cyber-capable models, and the hidden costs of agentic workflows—token budgets, security incidents, and comprehension debt—arguing that organizational redesign, not model capability, is now the real constraint.
The Daily Agentic AI Podcast - 2026-10-02
A large-scale dataset of real coding agent sessions reveals that only 59% of agent-written code gets committed, with frequent user pushback and more security vulnerabilities than human-written code, while separate studies show agents misspecify environments and generate library errors. Other key topics include the Pi 1.0 and Pi Durable harness releases, Claude Code's new mod system with Token Weather and Blast Radius features, model-harness interaction research showing rankings reverse across scaffolds, and new decision models from Cloudflare (Clef/Clef-flash), AWS Strands, and Perplexity. The episode also covers safety persistence in recursive self-improvement agents, the Misfit-Governed Development design theory, and a study of orchestration and tool integration issues in open-source multi-agent systems.
The Daily Agentic AI Podcast - 2026-10-01
Google DeepMind unveiled Gemini 4 Argon with a one million token output limit, strong benchmark results, and reported internal agent-driven kernel migrations, though access is restricted to trusted defenders with no public date. Other releases include Upstage's efficient Solar Mini 4, Ant Group's Ling 3.1 Flash, and growing Western adoption of Chinese cache optimizations. Research highlights span reliability and efficiency interventions (HiSentinel, HERO), whole-repository coding benchmarks (E2E-SWE, Zero2Repo, DoGBench, EngramBench), OpenCollab's multi-agent framework, specification preambles reducing defects, PatchHolmes patch retrieval, Aletheia permission testing, Runtime Assurance Contracts, verification-failure reuse, self-spec code generation, diffusion model code editing, and ARCHER compliance checks, alongside industry adoption barriers, Magnitude's inference engine, a Claude developer hub, rogue agent worm risks, OpenClaw Enterprise, Context Language Models, Perplexity's contextual embeddings, TomasuLLM out-of-order execution, Google's Agent Anomaly Detection, Mercury Voice, and Adaptive-GEPA routing.
The Daily Agentic AI Podcast - 2026-09-30
OpenAI launched Dots, always-on persistent agents powered by GPT-6 Astra that run on dedicated cloud computers with browser access and thousands of app integrations, priced for Pro and Business users with a safety monitor and human-review caveats. OpenAI also released GPT-6.1 Sol, a near-Astra-level agentic coding and computer-use model at roughly a fifth of the price with cached-input discounts, plus the Decisions API powered by GPT-6 Luna for real-time classification and routing. Other discussions covered Microsoft's AI Software Factory paper reporting 3x engineering efficiency on data systems, Assay's content-addressed evidence graphs binding agent claims to code hashes, a study finding AGENTS.md-style repository context files don't improve task success and raise inference cost over 20%, an executable-contract audit of tool-using agent benchmarks exposing benchmarks that don't measure what they claim, the Codex Security Cloud upgrade with cyber-capable models, Perplexity's Rust-based Photon retrieval engine with sub-100ms latency and Fast Search trade-offs, engineering patterns from Google's AI Agents Challenge, continual search for root-cause attribution in agent failures, XRepoSkill's transferable skills for SWE agents, and MemDocAgent's memory-guided repository documentation.
The Daily Agentic AI Podcast - 2026-09-29
Anthropic released Claude Sonnet 5.5, a faster and cheaper mid-tier model that reportedly scores over 70% on Anthropic's agentic coding benchmark versus roughly 10% for its predecessor, though it ships with stricter cyber safeguards that have drawn community complaints. OpenAI reopened the $200 ChatGPT Pro plan to new subscribers while halving the usage allowance, and Artificial Analysis launched a Cyber Index for evaluating agents on finding and fixing enterprise vulnerabilities. A case study showed off-the-shelf coding agents acting as goal-persistent performance engineers—autonomously removing PyTorch dependencies from a CUDA nearest-neighbor implementation—while other research covered the Say, Do, Understand workflow for developer understanding, SWE-Adept's two-agent localization/resolution framework, and lightweight governance approaches.
Een wekelijkse AI-gegenereerde podcast over het JVM-ecosysteem — Java, Kotlin, frameworks en meer. Beschikbaar als RSS feed.
De originele Sourcelabs Podcast — gesprekken over software engineering, teamdynamiek en het vak. Momenteel op pauze.