Ga naar de inhoud
PodcastsTechnologieAgentic Conversations (formally mlops.community)

Agentic Conversations (formally mlops.community)

Demetrios
Agentic Conversations (formally mlops.community)
Nieuwste aflevering

556 afleveringen

  • Agentic Conversations (formally mlops.community)

    The Caveman Prompting Challenge

    01-10-2026 | 44 Min.
    Caveman prompting has one rule: why use many words when few do the trick? It saves tokens on the way in and on the way out. Push it too far, though, and the output falls apart. So how far is too far? Nobody has benchmarked it yet, and that question opens our conversation with James Barney, Head of Forward Labs at MetLife.

    James spends his days connecting new AI capabilities to old business problems across dozens of regulatory regimes, and he still finds time to push code. He explains how the FinOps Foundation's AI working group took on the most basic question: which model for which workload, and why the answer always comes down to cost, speed, and accuracy. We get into Anthropic's launch pricing for Fable, why a million tokens is easy to price and hard to explain, and why every stakeholder eventually tells you what they really care about once you name the wrong North Star.

    From there it gets practical. Start with the smartest model, then step down and add harness until quality holds. Treat exploration tokens like local builds and production tokens like pipelines. Govern agents the way you govern people, with proactive blocks, reactive checks, and policy as code an agent can actually read.

    We close on a bigger shift. When a chat window can pull from every dashboard at once, do we still need dashboards? James thinks mostly not, with one catch he calls the latent shopper problem: some insights only come from browsing data you did not know to ask about.

    Demetrios Brinkmann: https://www.linkedin.com/in/dpbrinkm
    James Barney: https://www.linkedin.com/in/james-barney

    Timestamps:
    [00:00] Cold open
    [01:03] Meet James from MetLife
    [01:11] What AI enablement means at a global insurer
    [03:27] Inside the FinOps Foundation AI working group
    [05:03] The caveman skill challenge
    [06:37] Why we need a CaveBench
    [07:43] Anthropic's Fable launch pricing
    [09:34] Planning for surprise model releases
    [11:45] Measuring AI value beyond cost
    [12:47] Finding your unit metric
    [14:15] Tying token spend to business outcomes
    [16:35] Explore first, then optimize
    [18:12] AI that tunes itself
    [20:13] Do R&D tokens count
    [23:24] When the experiment becomes the product
    [25:38] Personal agents and daily briefings
    [27:39] Governing agents that run just because they can
    [30:37] The layers of AI governance
    [31:56] Proactive and reactive guardrails
    [32:59] Policy as code for agents
    [34:41] Stop the click ops
    [35:31] Is the modern UI obsolete
    [36:36] MCP apps and chat as the new browser
    [38:07] How AI gathers data differently than humans
    [40:37] The latent shopper problem
    [42:12] Staying close to your data
  • Agentic Conversations (formally mlops.community)

    AWS Has 16,000 APIs. Can MCP Handle It?

    28-09-2026 | 41 Min.
    What happens after you’ve built your first MCP server and actually have to make it work in the real world?

    In this episode, James Ward dives into the more advanced side of MCP: observability, evals, tool design, code mode, authentication, and the challenges that appear once agents start using your server at scale.

    We also get into how AWS thinks about MCP across roughly 16,000 APIs, why inefficient tool design often gets blamed on MCP itself, and whether the future could involve more constrained, human-reviewable alternatives to full code mode.

    Along the way, James shares a great example of an AWS documentation change that accidentally triggered prompt injection warnings from agents, showing just how complicated testing across different models and harnesses is becoming.

    If you’re already building with MCP and want to understand what comes after the “hello world” stage, this one goes deep.

    Timestamps:
    [00:00] “This Code Is Gobbledygook”
    [00:44] What Happens After You Build an MCP Server?
    [02:11] The MCP Patterns You Actually Need in Production
    [03:48] Why Your Agent Is Making Too Many Tool Calls
    [05:29] How to Make MCP Use Fewer Tokens
    [08:44] Is MCP Actually Inefficient?
    [10:11] “We Built a Lot of Pretty Crappy MCP Servers”
    [11:25] The MCP 2.0 Migration Problem
    [14:02] Why AWS Is Rethinking Code Mode
    [15:35] The Problem With Letting Agents Write Python
    [19:31] How Do You Measure Agent Experience?
    [22:11] Finding Out Why an Agent Failed
    [23:42] Hundreds of Evals for Four Cents
    [24:36] The Evaluation Matrix Gets Massive
    [26:28] AWS Accidentally Triggered a Prompt Injection Warning
    [30:16] Should MCP Servers Expose Only Five Tools?
    [31:02] AWS Has 16,000 APIs. Now What?
    [34:09] One Super-Agent or Thousands of Specialized Agents?
    [36:55] The MCP Authentication Problem
    [40:15] What’s Coming at AgentCon
  • Agentic Conversations (formally mlops.community)

    Skills over MCP on the streets of Tokyo

    24-09-2026 | 32 Min.
    Tool descriptions tell an agent what a tool does. They don't tell it how to use five tools together, in the right order, following your conventions. That gap is where this conversation lives.
    Filmed at AGNTCon + MCPCon in Tokyo with Ola Hungerford, Principal Engineer for AI Enablement at Nordstrom and a maintainer of the Model Context Protocol, who spent the last several months turning a pattern everyone was quietly reinventing into an actual MCP extension.
    Ola walks through what skills over MCP really means: the server stops being a pile of tools and becomes a distribution channel, handing the agent the instructions, workflows and knowledge it needs only at the moment it needs them. She explains why server instructions weren't enough, how progressive discovery keeps context from exploding, and why the same mechanism works for memory and preferences even when no tools are involved.Then it gets into the harder parts. What belongs in the MCP spec versus the agent skills spec. Why passing custom front matter through opens a rug pull and prompt injection surface nobody wanted. Where skills start to look like sub-agents, and why there's still no standard way to declare which servers a skill depends on. And the honest problem underneath all of it: how do you standardize something while everyone is still finding out what it's actually for, without breaking a hundred things the next time you change your mind?
    Timestamps:[0:00] Intro[0:21] AI enablement at Nordstrom[0:33] What skills over MCP actually is[1:29] The MCP server as a distribution channel[2:01] Server instructions versus skills[3:11] Distributing knowledge and memory[4:23] Progressive discovery explained[5:23] Where the idea came from[6:59] From draft to official extension[8:19] What early adopters changed[8:51] Front matter and custom metadata[9:57] Rug pulls and prompt injection risk[10:57] Will any of this get standardized[12:10] Marrying two very different specs[13:03] Skills as personas and sub-agents[13:53] The missing dependency standard[15:15] How the extension actually works[16:19] What harnesses still need to support[16:59] Consent and skill integrity[17:56] Where skills over MCP goes next[18:43] Why cramming 200 tools fails[19:41] Standardizing before you know the answer[22:18] Is git the wrong tool for agents[23:19] Picking tools for the actual persona[23:51] Trying to be less productive[25:48] The anxiety of idle agents[27:15] Why she keeps a robot on her desk[28:22] If the agent feels the friction, does it matter[29:41] Efficiency, waste, and caring enough[30:52] Letting an agent debug for you[32:03] Choosing your rabbit hole
  • Agentic Conversations (formally mlops.community)

    Walking Tokyo Talking Agent Protocols

    18-09-2026 | 28 Min.
    Two people, a wrong turn into a back alley, a community garden, and about thirty minutes of arguing about protocols on the streets of Tokyo.
    The guest is Angie Jones, VP of Developer Experience at the Agentic AI Foundation, fresh off launching AGNTCon + MCPCon in China before the Tokyo stop. She opens with what she learned there: a mobile-first, super-app world where the integration problem most of us obsess over barely exists, where every conversation about agents is really a conversation about the model, and where companies are now reaching for MCP and A2A precisely because they want to operate outside that ecosystem.The bulk of it is WebMCP - a protocol with a confusing name and, until recently, almost no attention. The pitch: put tool calling in the page itself, so your agent works inside your logged-in session with only the tools relevant to the page you're on, instead of screenshotting an anonymous browser and burning tokens guessing at the accessibility tree. Angie explains why it went from ignored to urgent the moment agentic browsing got good, and why the fix for computer use being slow and hijacking your machine might be a standard rather than a better model.
    It closes on agent-to-agent: whether anyone actually wants a marketplace of thousands of agents, or whether the real value is the one agent that has access you'll never get. Plus a well-earned complaint about three-letter acronyms and why researchers are still the only people naming things well.
    Timestamps:[0:00] Intro[0:59] Launching the conference in China[1:34] What North America gets wrong about agents[2:23] Super apps versus endless integrations[3:14] What happens when they expand beyond China[3:39] Tencent and A2A in production[4:34] A model-first country[5:56] Chinese coding agents and harnesses[6:59] Tokyo and the conference world tour[7:26] What WebMCP actually is[8:56] Why it has nothing to do with MCP[9:20] Page-level tools and your logged-in session[10:36] Why WebMCP sat unnoticed for months[11:25] Token efficiency and reliability[12:23] The moment computer use got good[13:15] Two real grievances with computer use[14:06] Collaborating instead of surrendering your screen[15:24] A short detour into Tokyo signage[16:15] Why web developers should be excited[17:04] Agents and the loss of first-party data[18:27] Why an agent cannot just buy something[19:23] Inside the agentic commerce working group[20:18] Upsells recommenders and an agent that ignores them[23:27] The commerce protocols to watch[24:19] Why A2A is next[25:31] Publishing your agent as a service[27:13] The case against agent marketplaces[27:55] Why access beats capability[29:31] Google's protocol land grab[30:45] Bring back the cool names[31:39] Amsterdam, San Jose, and what comes next
  • Agentic Conversations (formally mlops.community)

    Why Cost Per Million Tokens Is A Useless KPI?

    14-09-2026 | 38 Min.
    A year ago, Palo Alto Networks built dashboards to track AI spend. Today those dashboards are useless, and the team that built them thinks that's the whole story.
    Recorded at FinOps X in San Diego, this conversation brings together Abhinav Lad, who leads cloud and AI finance at Palo Alto Networks, and Kuntal Patel, who runs the cloud engineering function behind it. They explain what happened when agents entered the picture, and AI stopped behaving like a service anyone could forecast.
    The short version: consumption went from linear to exponential almost overnight. Agents are goal-oriented rather than task-oriented, so they plan, call tools, verify, fail, retry, and keep looping until they hit the outcome, and every iteration is billable.
    So how do you run finance on top of that? Abhinav and Kuntal walk through the metrics that replaced their old forecasts: adoption rate, cost per user, AI as a percentage of revenue - and the budget limits that let engineering leaders choose between the newest model and a longer runway. They get into the open question of whether a cheaper model saves money or just burns more tokens thinking. They explain why an AI gateway became the control plane for cost and security at the same time, why retry caps belong in the design phase instead of the postmortem, and how FinOps starts to resemble product QA once the bill becomes the clearest signal that something is broken.
    They close on a warning worth sitting with: cost per million tokens is a number that means almost nothing on its own, and a value story built on it will point you somewhere you don't want to go.

    Palo Alto Networks: https://www.paloaltonetworks.com

    Abhinav Lad: https://www.linkedin.com/in/abhinav-lad
    Kuntal Patel: https://www.linkedin.com/in/kuntalpatel35
    Alex Salkever: https://www.linkedin.com/in/alexsalkever

    Timestamps:
    [0:00] Intro
    [1:00] Who runs FinOps for AI at Palo Alto Networks
    [2:10] Last year's AI dashboards are already useless
    [4:26] Agents turned linear forecasts exponential
    [7:27] Three traits that make agents expensive
    [8:34] The hidden bill: RAG, vectors and egress
    [9:16] Cost per user and adoption rate
    [11:21] Giving engineering leaders a budget
    [12:07] Using DORA metrics to prove value
    [13:53] Where DORA stops fitting AI
    [16:20] Does the cheaper model actually save money
    [17:57] Why you need an AI gateway
    [20:05] Inside Prisma AIRS
    [21:00] Three cost models for three use cases
    [22:52] Forecasting lessons from Electronic Arts
    [24:03] Runaway agents and endless loops
    [25:59] Capping retries before they burn cash
    [28:06] Writing cost policy at design time
    [29:01] When FinOps becomes product QA
    [32:17] Explaining AI spend to the C-suite
    [34:51] Valuing AI beyond engineering
    [37:04] Crawl, walk, run: where they are today
    [38:20] Why cost per million tokens is meaningless
    [39:26] Closing thoughts
Meer Technologie podcasts
Over Agentic Conversations (formally mlops.community)
Relaxed conversations and technical deep dives around AI Agents. This Show is brought to you by the Agentic AI Foundation where the leading agentic open-source projects like MCP, Agents.md, and Goose live. See more at aaif.io
Podcast website

Luister naar Agentic Conversations (formally mlops.community), Tech Update | BNR en vele andere podcasts van over de hele wereld met de radio.net-app

Ontvang de gratis radio.net app

  • Zenders en podcasts om te bookmarken
  • Streamen via Wi-Fi of Bluetooth
  • Ondersteunt Carplay & Android Auto
  • Veel andere app-functies
Agentic Conversations (formally mlops.community): Podcasts in familie