How To Run A 90-Day AI Pilot On A Marketing Team
A 90-day AI pilot needs workflow selection, baseline measurement, controlled rollout, and a final team decision.
Short answer
A 90-day AI pilot on a marketing team has three phases: workflow selection and baseline measurement (weeks 1–4), controlled rollout with daily output review (weeks 5–8), and measurement and team decision (weeks 9–12).
A 90-day AI pilot is not a proof-of-concept. It is a structured decision process. By the end, you should know whether a specific AI workflow is worth operating at scale, whether your team can run it consistently, and what it actually costs in time and money.
Most pilots fail not because the AI underperforms, but because there was no clear decision criteria at the start. Ninety days pass, the team has impressions but no data, and the question "should we keep doing this?" gets answered by whoever is most enthusiastic in the room.
Here is a structure that produces a real answer.
The three-phase structure
| Phase | Weeks | What you are checking |
|---|---|---|
| Select and baseline | 1–4 | Which workflows to test; what they currently cost; what good output looks like |
| Controlled rollout | 5–8 | Whether the AI-assisted workflow produces output at or above baseline quality, consistently |
| Measure and decide | 9–12 | Whether the improvement holds; whether to expand, pause, or kill |
Each phase has a gate. You do not move forward unless the gate criteria are met. This is the part most pilots skip, and it is why they drift.
Phase 1: Select and baseline (weeks 1–4)
Pick one or two workflows, not ten. The instinct is to pilot everything at once. Resist it. You cannot measure ten things carefully. You can measure two things well.
Good pilot workflows share three traits: they run on a regular cadence (weekly or more), they produce a reviewable output (a document, a report, a draft), and they have a clear definition of "done well." If you cannot describe what a good output looks like before the pilot starts, the pilot will not teach you much.
During weeks 1–4, run the workflow manually and document it. How long does each step take? What does the output look like when the person doing it has a good week versus a bad week? What is your revision rate?
This is your baseline. Write it down. Leadership will ask for it in week 10.
Gate to Phase 2: You have a documented baseline for at least one workflow, with time-per-run, output quality criteria, and revision rate.
Phase 2: Controlled rollout (weeks 5–8)
Introduce AI assistance for one person or one sub-team, not everyone at once. You want to compare AI-assisted output to manual output under similar conditions — same week, same brief, similar scope.
Set a daily output review rhythm, even if it takes five minutes. Someone needs to look at each AI-assisted output and mark it: passed quality bar, needed minor revision, needed major revision, rejected. You do not have to do this forever. You do it now so you have data.
Track deviations. When the AI produces something unexpected or wrong, document what triggered it. Inconsistent input data is the most common culprit. If you find a pattern, fix the input structure before you blame the model.
Gate to Phase 3: AI-assisted outputs are passing your quality bar at the same rate or better than manual, for at least four consecutive weeks.
Phase 3: Measure and decide (weeks 9–12)
Pull the numbers together. Time saved per run. Quality pass rate. Cost per qualified output. Compare to baseline.
This is also when you run your "what did we learn" review with the team. Not "was this good or bad" — that is impressionistic. Ask: what did we expect to happen, what actually happened, and what explains the difference?
At week 12, you make one of three decisions:
- Expand: The workflow is operating above baseline on all three metrics. Roll it out to the full team.
- Pause: One or two metrics are marginal. Redesign the input structure or the review checkpoint, then re-pilot for four more weeks.
- Kill: Quality is below baseline after phase two. The workflow is not a fit for AI assistance at this stage. Document why and move to a different candidate.
"Kill" is not failure. It is data. Teams that cannot make a kill decision on a pilot will spend two years half-committed to a workflow that doesn't work.
How to get leadership buy-in before you start
Present the pilot as a decision framework, not a technology experiment. Leadership is not nervous about AI — they are nervous about spending money and time without a clear outcome.
Your ask is simple: four people, two workflows, 90 days, and one clear yes/no decision at the end. Show the three-phase table. Name the gate criteria. Tell them what success looks like and what the kill threshold is.
When leadership can see the decision structure, they can sponsor it. When it looks like "we are going to try some AI things," it is hard to fund or prioritize.
How Prova is structured the same way
The Prova program runs on sprint cycles with explicit review gates. You build something, it gets reviewed against stated criteria, and you either pass or revise. No drift, no vague impressions.
That rhythm — baseline, build, review, decide — is the same logic as the 90-day pilot. The sprint structure just compresses the timeline so you get faster feedback loops. If you want the underlying strategy for how to structure a larger rollout, the 90-Day AI Rollout Plan post covers that. This post is the execution layer: what to check at each gate, and what a real decision looks like at week 12.

