Ga naar de inhoud
PodcastsTechnologieProduct Impact Podcast | Secrets to unlocking the value of AI

Product Impact Podcast | Secrets to unlocking the value of AI

Presented by PH1
Product Impact Podcast | Secrets to unlocking the value of AI
Nieuwste aflevering

70 afleveringen

  • Product Impact Podcast | Secrets to unlocking the value of AI

    17. Your AI Product is Failing — Microsoft's UXR Team Knows Why

    20-07-2026 | 42 Min.
    In our last episode, Stanford NLP researcher Dr. Moritz Sudhof showed that 79% of AI product failures are invisible — they don't fire alerts, don't surface in telemetry, and don't get flagged by users, and they quietly erode trust and accelerate churn. This episode is the operational follow-up: the team that built a method for catching exactly that class of failure.

    Microsoft's Copilot UX research team spent a year running evaluations on real conversations — every archetype, every industry, every use case — and found that more than half of their quality failures weren't in any eval they were running. At the world's most widely deployed enterprise AI product, with sophisticated engineering and testing infrastructure, standard evals were still missing the majority of what users actually experienced as failure.
    That finding isn't limited to Copilot's scale. It's a structural gap in how the industry evaluates AI quality — and if you're running automated evals and calling that sufficient, the gap in your own product is almost certainly larger than you know.
    In this episode we cover:
    Token usage and adoption tell you if your AI is being used — not whether it's actually working for anyone.
    Users bring real prompts, test one model fully, then compare — that's what produces honest signal at scale.
    More than half of Copilot's loss patterns — user-driven gaps in model behavior — weren't in any existing eval.
    LLM judges get you to baseline quality. Users catch what automated testing structurally cannot.
    The flywheel: UXR evals → loss pattern taxonomy → log inspection → prompt changes → retention gains.
    This team started with 10 users and one comparative question. Signal strong enough to scale to an entire org.

    "More than half of the loss patterns that we've detected were not things that we were measuring in our evals." — Wendy Wang

    About the team: This work was developed by Christopher Monnier, Wendy Wang, and Chuck Kwong, UX researchers on the Microsoft Copilot team. Together, they built and continue to refine an interactive evaluation method that brings real user tasks, side-by-side product comparisons, quantitative results, and qualitative feedback into one process. The team uses this work to identify where Copilot succeeds and where people run into issues, understand the reasons behind user preferences, and turn the findings into clear opportunities for product and prompt teams to improve the experience.
    Christopher Monnier on LinkedIn: https://www.linkedin.com/in/christophermonnier/
    Chuck Kwong on LinkedIn: https://www.linkedin.com/in/charleskwong/
    Wendy Wang on LinkedIn: https://www.linkedin.com/in/wendy-wang-mertensmeyer-8386b018/
    Microsoft Copilot: https://www.microsoft.com/en-us/microsoft-copilot
    ..................
    If you found this episode useful, please like, share, and send it to anyone on your team who'd find it helpful.
    We built https://productimpactpod.com to be your AI product insights and strategic playbook hub. Check it out.

    Hosted by:
    ➜ Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
    ➜ Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/

    Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com

    This episode was brought to you by:
    ➜ PH1 (https://ph1.ca) — a strategy & research consultancy specialized in pinpointing how to best leverage AI and improve the impact of your AI product
    ➜ AI Value Acceleration (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.
  • Product Impact Podcast | Secrets to unlocking the value of AI

    16: Invisible Failures: Stanford's Research on 100K AI Conversations — Moritz Sudhof, Bigspin AI

    08-07-2026 | 50 Min.
    KPMG pulled a report this year after its team accepted hallucinated information without question. Latham & Watkins submitted a court filing built on fabricated legal citations an AI invented, delivered with total confidence and perfect formatting. McDonald's AI ordering chatbot became a public failure for the same underlying reason: the system looked like it was working. Dr. Moritz Sudhof, CEO of Bigspin AI, analyzed 100,000 real conversations between users and live AI systems with Stanford NLP Group's Chris Potts and found why: 79% of AI failures are invisible to standard monitoring because they are behavioral failures, not technical ones.

    Sudhof built this research on hard-won experience. As VP of AI at BetterUp, he shipped an AI coach that beat ChatGPT on every expert coaching benchmark and still lost users — until his team changed nothing but how the AI introduced itself, and outcomes doubled. That gap between what an eval measures and what actually happens in the room with a user became his research question, and then his company.

    In this episode:
    The Confidence Trap: AI states something false, dressed in precise numbers and total confidence.
    Silent Mismatch: when the AI can't finish a task, it quietly answers a different question instead.
    The Drift and the Death Spiral: how a long conversation loses the plot and burns the user's patience.
    Models are post-trained to answer, not clarify — a default that gets worse as models get more capable.
    The Paradox of AI Fluency: the users who push back hardest also hit the most failures, and succeed most.
    Evals catch what you already know to test for. Reading real transcripts catches what you don't.

    "The behavioral layer is where most of the damage is actually happening." — Moritz Sudhof
    "The more specific and helpful it often gets, the more fake it often is." — Moritz Sudhof, on the Confidence Trap

    About Moritz: Shipped conversational AI to hundreds of organizations as VP of AI at BetterUp. Lived the problem of knowing something's wrong but not what to fix.
    Guest resources:
    LinkedIn: https://www.linkedin.com/in/sudhof/
    X: https://x.com/mmooritz
    Personal site: https://msudhof.com/
    Bigspin website: https://bigspin.ai
    Key research:
    Invisible Failures in Human–AI Interactions: https://arxiv.org/abs/2603.15423
    A paradox of AI fluency: https://arxiv.org/abs/2604.25905

    We built productimpactpod.com to be your AI product insights and strategic playbook hub. Check it out.
    Thank you for listening to the Product Impact Podcast — if you have feedback, guest recommendations, or want to chat — contact us.

    Hosted by:
    Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
    Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/

    Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com

    This episode was brought to you by:
    PH1 (https://ph1.ca) — a strategy & research consultancy specialized in delivering evidence about the highest value use cases and customer profiles. AI Value Acceleration
    (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.
  • Product Impact Podcast | Secrets to unlocking the value of AI

    15. Playbook for Increasing AI Adoption & Value Creation

    25-06-2026 | 26 Min.
    Four data reports from 2026 tell a consistent story, and none of it matches the adoption narrative. Writer surveyed 2,400 global workers and C-suite leaders: 97% of executives deployed AI agents in the past twelve months, 29% reported significant ROI. Glean's Work AI Index found a name for what most knowledge workers are actually experiencing: botsitting — spending more time supervising and correcting AI than gaining anything back. Section's biannual proficiency survey: 67% of workers use AI weekly, 5.5% are proficient enough to generate consistent value, and 79% of managers haven't demonstrated their own AI use to their team in the past month. Token consumption per organization grew roughly 320 times in twelve months while the share of organizations reporting significant ROI stayed at 29%.
    Brittany Hobbs and Arpy Dragffy work through what's causing the gap, why the teams trying hardest to close it keep hitting structural walls, and what it takes to move from measuring adoption to generating defensible value — for individual contributors, for teams, and for the organizations responsible for this investment.

    What you'll learn:
    OpenAI, Writer, Glean, Section: four reports, one consistent signal — the adoption story hides the value failure.
    Glean 2026: botsitting is the dominant AI experience. More knowledge workers are losing time to AI than gaining it.
    67% of workers use AI weekly. Only 5.5% are proficient. The problem isn't more training days. It's the model of change.
    Four years of measuring seats over outcomes has left AI leaders unable to defend their budgets. The window is closing.
    Salesforce agreed to acquire Fin for $3.6B. What they built before that exit is the lesson most orgs are ignoring.
    Boris Cherny no longer prompts — he builds loops. What that means for every team not yet running autonomous evaluation.

    Articles referenced in this episode:
    97% of Executives Deployed AI Agents. Only 29% See ROI. — Brittany's breakdown of the Writer 2026 survey and the 68-point deployment-to-value gap
    The 10% Problem: AI's Value Gap Is Wider Than Anyone Is Admitting — Why AI value is concentrating at the top and what it means for the rest of the organization
    WTF is an AI-native org anyways? Let's compare Airbnb & Meta's opposing plans. — The competing models for AI-native organization design
    OpenAI & Anthropic are charging us way more than we need — Arpy on token economics, model selection, and the cost side of AI value creation

    We built productimpactpod.com to be your AI product strategy and AI product news hub. Check it out.
    Thank you for listening to the Product Impact Podcast — if you have feedback, guest recommendations, or want to chat — contact us.
    Hosted by:
    Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
    Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/
    Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com

    This episode was brought to you by: PH1 (https://ph1.ca) — an AI strategy consultancy specialized in improving the measurable success of AI products. AI Value Acceleration
    (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.
  • Product Impact Podcast | Secrets to unlocking the value of AI

    14: AI Adoption is the Problem Everyone is Desperate to Solve — Dr. Molly Sands, Atlassian

    16-06-2026 | 31 Min.
    Six months of research into the world's leading AI-powered organizations reveals a consistent split: a handful of people are seeing 10x or 20x gains, most are seeing some movement, and a significant portion of the workforce is drowning in forced change — trying to keep up with tools and mandates while watching colleagues get laid off. The organizations pulling ahead aren't pushing harder. They're leading by example, building cultures where struggling out loud is allowed, and being honest about where they actually are in the AI journey. The ones still stuck are running on fear-based incentives, measuring adoption instead of value, and missing the governance infrastructure — no Chief AI Officer, no clear policies, no connective tissue between independent AI experiments.
    Atlassian's 2026 State of Teams report puts numbers to the pattern. Twelve thousand knowledge workers, 170 Fortune 100 executives, and a headline finding: the Fortune 500 is losing $160 billion a year to what Atlassian calls the AI fragmentation tax — the cost of everyone moving fast in different directions. Dr. Molly Sands leads the Teamwork Lab at Atlassian, where behavioral scientists study how teams work and what separates high-performing ones from the rest. Her team found that organizations seeing real AI ROI moved to team-level AI thinking first — redesigning shared workflows instead of letting individuals invent their own, creating AI working agreements that give people clarity instead of anxiety, and breaking down knowledge silos rather than restructuring org charts. Information flow turned out to matter more than reporting structure.
    The episode also gets into what the research shows about junior employees (they're more comfortable than their managers), whether 2026 is actually the year of the agent (it isn't — not yet, not at scale), and what it's going to take to stay relevant once simply adopting AI stops being enough.

    Why AI adoption is still uneven — and what "drowning in forced change" actually looks like inside organizations
    Why the governance gap — no CAIO, no policies, no connective tissue — is the real reason AI experiments don't compound
    Why the Fortune 500 is losing $160 billion a year to coordination chaos, and why better tools won't close that gap
    Why team-level AI thinking drives faster ROI than individual adoption programs or usage mandates
    What AI working agreements are, what Atlassian's research found when teams used them, and how to run one
    Why most companies are nowhere near the orchestration level — and what the AI maturity curve actually looks like from the inside

    "Just saying 'go off and try it' can actually feel really hard. The more clarity around what you have access to and how you can use it — the better the teams tend to do." — Dr. Molly Sands, Atlassian
    ----
    If you found this episode useful, please like, share, and send it to anyone on your team who'd find it helpful.
    We built https://productimpactpod.com to be your AI product strategy and AI product news hub. Check it out.

    Hosted by:
    Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
    Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/
    Featured guest:
    Dr. Molly Sands — https://www.linkedin.com/in/mollysands
    Atlassian 2026 State of Teams Report — https://www.atlassian.com/blog/teamwork

    Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com

    This episode was brought to you by:
    PH1 (https://ph1.ca) — an AI strategy consultancy specialized in improving the measurable success of AI products.
    AI Value Acceleration (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.
  • Product Impact Podcast | Secrets to unlocking the value of AI

    13. Why Managing AI Agents Is More Like Supervising Labor Than Using a Tool [Jonathan Su, Procurify]

    08-06-2026 | 30 Min.
    Managing an AI agent isn't using a tool — it's supervising labor. Most companies skipped that step. In Procurify's recent survey of finance leaders, 35% said trust — not model capability — is the single biggest factor in whether their organization can actually deploy agents. The teams already shipping report 63% ROI from time savings and 60% from improved data accuracy, but only after they did the unglamorous work first: defined the operating model, baked in governance and audit trails, and consolidated their data into a single source of truth. Frontier models keep commoditizing generic intelligence. The value is moving up the stack — to the workflow, the context, and the data your company actually runs on.
    Procurement has sat in the middle of every enterprise's audit trail for decades — budgets, contracts, suppliers, approvals, compliance, payments. It's the use case AI vendors have been quietly building toward, because if you can make procurement feel less clunky, you've solved governance for the rest of the business. We sat down with Procurify's Chief Product & Technology Officer Jonathan Su to understand what an AI-native operating model actually looks like, why production-grade is now ten times harder than prototype, and what shifts when the bottleneck in your team moves from execution to judgment.

    In this episode:
    Why 35% of finance leaders say trust — not model capability — is the biggest factor in whether agents actually ship
    The operating model most companies skip: governance, audit trail, single source of truth — before the agent touches work
    What AI ROI actually looks like — 63% time savings, 60% better data accuracy, plus the business KPIs that prove it
    Why value is moving up the stack as frontier models commoditize generic intelligence — workflow, context, data, distribution
    How procurement teams redesign workflows around agents instead of tacking AI on top of an already broken process
    The hire that beats 20 years of experience: grit, taste, judgment, and the ability to learn in 4-month cycles

    "Managing an agent is more than just using a tool. It's sort of like supervising labor." — Jonathan, Procurify
    "The cost of producing something is dramatically lower, but the bottleneck shifts to judgment, craftsmanship, and taste. Just because you could do something doesn't mean you should." — Jonathan, Procurify

    We built productimpactpod.com to be your AI product insights and strategic playbook hub. Check it out.
    Thank you for listening to the Product Impact Podcast — if you have feedback, guest recommendations, or want to chat — contact us.

    About Jonathan: Jonathan is Chief Product Officer at Procurify, where he leads product strategy and AI initiatives across the company's spend management platform. He has spent his career in payments, fintech, and enterprise software, and now leads Procurify's transition to an AI-native product organization. Procurify serves finance teams managing budgets, approvals, invoicing, and payments — the workflows where governance and AI agents have to coexist. 
    Procurify: ⁠https://www.procurify.com⁠
    Jonathan on LinkedIn: ⁠https://www.linkedin.com/in/jonathanhaosu⁠

    Hosted by:
    Arpy Dragffy Guerrero — https://www.linkedin.com/in/adragffy/
    Brittany Hobbs — https://www.linkedin.com/in/brittanyhobbs/

    Go to Substack to get AI strategy frameworks, news, and jobs: https://productimpactpod.substack.com

    This episode was brought to you by:
    PH1 (https://ph1.ca) — a strategy & research consultancy specialized in delivering evidence about the highest value use cases and customer profiles.
    AI Value Acceleration (https://aivalueacceleration.com) — The consultancy specialising in enterprise value creation. Make sure that your spending doesn't go to waste. Find out exactly where the value creation of adopting AI products stalls.
Meer Technologie podcasts
Over Product Impact Podcast | Secrets to unlocking the value of AI
No-nonsense advice and strategies from AI product leaders, designers, and researchers Learn how to overcome adoption barriers and scale impact across teams and customer bases. Our audience learns powerful insights that will shift how they think about and leverage AI. At the core is how to improve the UX of using AI and to enhance the quality and consistency of the products we depend on most for work. Resources and playbooks: https://productimpactpod.com Hosted by Arpy Dragffy Guerrero (PH1 — https://ph1.ca) and Brittany Hobbs (AI Value Acceleration — https://aivalueacceleration.com).
Podcast website

Luister naar Product Impact Podcast | Secrets to unlocking the value of AI, De Technoloog | BNR en vele andere podcasts van over de hele wereld met de radio.net-app

Ontvang de gratis radio.net app

  • Zenders en podcasts om te bookmarken
  • Streamen via Wi-Fi of Bluetooth
  • Ondersteunt Carplay & Android Auto
  • Veel andere app-functies