Michael Marriage
Back to blog

Your AI Pilot Shouldn’t Require a Fire Extinguisher

December 16, 2025

AIProduct Management

Lessons from running a real AI alpha/pilot that didn’t set anything on fire and helped me preserve what's left of my sanity.

If there’s one truth about launching AI features, it’s this: everyone wants the magic, no one wants the meltdown.

The hype is intoxicating. Executives want it, customers expect it, competitors brag about it, and your product roadmap suddenly reads like it was written during a tornado made entirely of sticky notes..

But here’s the uncomfortable part: AI is unpredictable, especially in early development. It is brilliant right up until the moment it enthusiastically hallucinates a fictional donor, rewrites someone’s job title, or recommends something that violates every best practice your company has ever stood for.

That’s why running an AI alpha or pilot is both a privilege and a responsibility. You’re not just testing features, you’re testing trust.

I recently ran both an AI Alpha and Pilot program at my company. We were building new intelligence-driven capabilities that would shift core workflows for customers and redefine how our platform interacts with their data. The stakes were high. And the biggest goal was simple: prove the value while protecting the trust we’ve worked years to earn.

Here’s what I learned; what worked, what mattered, and what I recommend for anyone about to shepherd an AI product through its earliest customer-facing stages.

Start With “Why Now?” Not “Why AI?”

Before you ever invite a customer into an alpha, you need a crisp answer to a deceptively simple question:

"Why should this exist right now?"

AI for AI’s sake is a fantastic way to confuse customers and irritate engineers. Instead, define:

  • The pain: What real customer problem requires intelligence, not just automation?
  • The opportunity: What outcome becomes possible with AI that wasn’t possible before?
  • The timing: Why test this now, not later?

In our case, the ‘why now’ was undeniable: customers had a lot of great data in their CRM, but lacked both the donor insight and next best action guidance needed to compete in a noisy fundraising world. Giving them clear, intelligence-driven direction is essential to helping them raise more. The pain was real, the value was provable, and the timing aligned with our broader product strategy.

A pilot without a strategic purpose is just a demo with extra steps.

Alpha ≠ Pilot. Treat Them Differently

Here’s a simple rule of thumb:

  • Alpha: messy, exploratory, feedback-heavy, internal or highly trusted customers only.
  • Pilot: more polished, measurable, broader but still limited, validating readiness for scale.

The biggest mistake teams make is treating alphas like pilots. That is, trying to make them too perfect.

Don’t. Do. This.

You want alphas to reveal the sharp edges. You want to learn where things break. You want honest feedback before marketing ever lays eyes on it.

When we ran our alpha, we deliberately framed it as:

“You’re co-building with us. We expect weird behavior, and we need your eyes on it.”

This framing creates safety. For your team and for your customers.

Design Your Alpha Group Like a Casting Call

Not every customer belongs in an alpha or pilot. In fact, many shouldn’t be anywhere near it. You want customers who are:

  • Technically comfortable with new workflows.
  • Experiment-friendly and not expecting perfection.
  • Collaborative and willing to provide structured feedback.
  • Invested in the outcome, not just the novelty.

You also want a mix:

  • Power users with deep workflow familiarity.
  • Customers that represent the “typical” user.
  • Customers that push edge cases.
  • A few who are skeptical enough to keep you honest. (But avoid complete AI detractors at this point.)

In our alpha, we handpicked customers we trusted, and who trusted us. It made all the difference.

Set Expectations Like Your Product Depends on It (Because It Does)

AI changes the contract between you and your users. You’re asking them to trust a system that will sometimes be wrong, occasionally be weird, and inevitably do things no one predicted.

So you need to be painfully clear about:

What it can do

(“This will generate a donor email draft based on your notes.”)

What it cannot yet do

(“It will not send that email, adjust donation data, or summon ancient spirits.”)

What risks exist

(“It may occasionally produce content that needs correction.”)

How to give feedback

(“There’s a button right here. Please click it generously.”)

How we are monitoring

(“We are reviewing outputs daily; nothing is going to production unattended.”)

Underset expectations. Overdeliver on communication.

Instrument Everything. Not Later, Now

If you wait to add instrumentation until after the pilot, congratulations: you have just run a feelings-based experiment. AI requires data about the data about the data.

Instrument for events and outcomes such as:

Regenerations – How often users ask the AI to try again, signaling dissatisfaction, uncertainty, or a mismatch between intent and output.

Edits – The degree to which users modify AI-generated content, revealing where the system falls short of accuracy, tone, or usefulness.

Overrides – Instances where users bypass or undo AI suggestions entirely, indicating low trust or poor alignment with user expectations.

Task completion rate – The percentage of AI interactions that successfully result in the intended outcome without manual rework.

Time saved – A measure of how much effort the AI actually removes from the workflow compared to doing the task manually.

User confusion moments – Points where users hesitate, abandon, or seek clarification, often signaling unclear UX or unexpected AI behavior.

Hallucination patterns – Recurring situations where the AI generates incorrect or fabricated information, helping identify systemic risk areas.

Guardrail triggers – Events where safety, permission, or policy constraints activate, revealing both protection effectiveness and friction points.

Drop-off points – Stages in the AI workflow where users disengage, highlighting where trust, clarity, or value breaks down.

This data tells you what no customer will ever articulate out loud.

Create a “Trust Buffer” Around AI Outputs

Users need to feel safe interacting with AI-generated content. That means:

  • No AI steps should be irreversible.
  • Users should always have an editing step.
  • Add visual cues (badges, context labels) to show what’s AI-generated.
  • Provide “Why am I seeing this?” explanations.
  • Store the original input for transparency.
  • Make overrides extremely easy.

We built in a multi-layer safety net during our pilot: AI suggestions were never final. It was always a draft, always editable, always inspectable. We also took an additional step of limiting bulk actions and record deletes to further safeguard customer data.

That design choice prevented multiple potential trust issues from ever surfacing.

Move Quickly, But Not Recklessly

AI development moves fast. AI deployment should not.

The temptation, especially after an exciting alpha, is to sprint to General Availability. But the responsible path is:

Validate value - Validate safety - Validate understanding - Validate consistency - Validate UX simplicity

If even one of these is missing, your AI is not ready.

A pilot isn’t a prelude to launch. It’s a filter for whether launch is earned.

**Feedback is Gold **

But Only If You Don’t Drown in It

AI pilots generate feedback like a toddler generates chaos, continuously and creatively.

Categorize it as follows:

Critical

Data leakage, RBAC issues, hallucinations, unsafe outputs, workflow blockers.

High-value

UX friction, clarity issues, misunderstood prompts.

Nice-to-have

Could it also…?”

(Yes, someday. No, not today.)

Emotional responses

I love this,” “I’m nervous,” “This is weird.

(All useful, none actionable alone.)

When we ran our pilot, we built weekly feedback debriefs with product, engineering, and CX. This cadence kept us grounded, aligned, and sane (well, some of us anyway.).

Measure Value Like Your Life Depends on It

AI pilots should produce quantified outcomes, not just vibes.

Examples include:

Time saved - This measures how much effort the AI removes from a workflow compared to doing the task manually. For example, a process that previously took 30 minutes is completed in under five with AI assistance.

Steps removed - This tracks how many manual actions or decisions the AI eliminates from an existing workflow. For example, an AI-generated donor summary replaces five separate clicks across multiple CRM screens.

**Errors prevented **- This measures the reduction in mistakes caused by manual data entry, interpretation, or decision-making. For example, AI validation catches incorrect donor segmentation before an outreach campaign is sent.

Content quality improvements - This evaluates whether AI-generated or AI-assisted content is clearer, more relevant, or more effective than human-only output. For example, donor emails generated by AI receive higher open and response rates than previous templates.

Faster completion - This captures how quickly users can finish a task end-to-end with AI support. For example, a volunteer follow-up workflow that once took hours is completed in minutes.

Higher throughput - This measures the increase in volume of work users can complete in the same amount of time. For example, a fundraiser can prepare outreach for 50 donors instead of 10 in a single afternoon.

Reduced cognitive load - This reflects how much mental effort the AI removes by simplifying decisions and surfacing next best actions. For example, instead of analyzing multiple reports, the user sees a clear recommendation on who to contact next.

**Increased user confidence **- This measures how comfortable users feel trusting the system’s guidance and acting on its recommendations. For example, users move from double-checking every AI suggestion to confidently executing recommended actions.

One of our pilot customers told us the new AI workflow saved them days of effort in their 2026 planning. That one sentence validated months of investment, and gave us momentum for broader rollout.

Communicate Progress Relentlessly

Silence kills trust. Updates build it. Share: Weekly improvements, known issues, fix timelines, new capabilities, insights learned, what’s coming next, what’s not coming next (equally important).

When customers know they’re part of the journey, not test subjects, they give better feedback and show more patience.

End the Pilot Gracefully

A good pilot has a beginning, middle, and end. Not an endless zigzag of “just a bit longer.” When wrapping up:

  • Share the results
  • Celebrate participant impact
  • Thank them meaningfully
  • Explain next steps
  • Set expectations for GA timeline

People love being part of something special. They hate feeling strung along or used.

Closing Thought

Running an AI alpha or pilot is one of the most delicate, high-stakes phases in building AI products. It requires humility, transparency, structure, and a deep respect for customer trust. Having recently navigated this process myself, from alpha to pilot to broader rollout, I can confidently say this: the quality of your early program shapes the quality of your entire AI strategy.

Run it thoughtfully. Run it responsibly.

Run it like trust is your most valuable feature. Because in AI it absolutely is.

Wishing you all the very best

Mike