Decision Guide · 5 minute read
How to Run AI Discovery: Baselines, Feasibility, and a Spec in Weeks
AI discovery is a short, structured phase before build that measures the baseline, defines what correct means with domain experts, tests feasibility by running candidate models against a first golden dataset, verifies data access, permissions, and integration paths, models cost at expected adoption, and produces a specification with acceptance criteria and a decision to proceed, narrow, or stop.
Most AI projects skip discovery or mistake a demo for it. They fund a build on an estimated benefit, discover in month two that the data is locked, and learn in month six that the accuracy users need was never reachable. Discovery is the short phase that answers whether and how before the expensive part: baseline measured, correctness defined, feasibility tested with evaluation, data and integration verified, cost modeled, specification written, decision made. This guide covers how to run it, drawing on FISTA Solutions' AI enablement practice. The specification it produces is in how to write an ai spec and the stage that follows in the ai pilot checklist.
What does discovery answer?
| Question | Activity | Deliverable |
|---|---|---|
| How big is the problem? | Baseline measurement from operational data | Baseline report |
| What does correct mean? | Working sessions with domain experts on real cases | Correctness definitions; first golden dataset |
| Is it feasible? | Candidate approaches evaluated against the dataset | Evaluation results by category |
| Is the data usable? | Access, permissions, quality verified with samples | Data readiness report |
| Can it be integrated? | Target systems, APIs, and permissions inventoried | Integration inventory |
| What will it cost? | Build and run cost modeled at expected adoption | Cost model |
| What could go wrong? | Risks identified with controls | Risk register |
| Should we proceed? | Evidence assembled | Specification and decision |
How do you measure the baseline?
From operational systems, not interviews: volumes, cycle times, cost per unit, error and rework rates, backlog, and satisfaction for the target process over a representative period. Segment by the categories the AI would handle differently. The baseline decides whether the problem is worth solving and becomes the reference for every later value claim. Case structure is in ai business case template.
How do you define correct?
In working sessions with the people who do the work today, using real cases: what a correct output is, what an acceptable error is, which categories matter, and what must never happen. Capture the results as labeled examples, which become the first golden dataset, and as written acceptance criteria by category. This is the step most often skipped and the one that determines whether the project can succeed. Dataset design is in what is a golden dataset.
How do you test feasibility?
Run candidate models and approaches, hosted and open, retrieval variants, prompt strategies, against the first golden dataset with calibrated grading, by category, and compare with the draft thresholds. The result shows which categories are reachable, which need narrowing, and what each approach costs per correct output. A demo shows possibility; evaluation shows reachability. Evaluation practice is in the ai evaluation checklist.
How do you verify data and integration?
Obtain sample data from each source with permissions confirmed, assess quality on the fields the system needs, and document lineage and residency constraints. Inventory the target systems, verify API access and rate limits, and identify where actions will need gates. Discovery that ends with samples in hand and API calls tested has removed the two most common causes of stalled builds. Readiness practice is in the ai data readiness checklist.
How do you model cost and risk?
Model build cost from the specification and run cost at expected adoption including model usage, human review, infrastructure, and maintenance, with routing and caching assumptions stated. Identify risks with controls and owners in a register. Both feed the decision and the staged funding plan. Budgeting is in the ai budget planning checklist and controls in how to de-risk an ai project.
What is the decision at the end?
Proceed to pilot with the specification and funding for the next stage; narrow scope to the categories evaluation showed reachable; or stop because the baseline is too small, the data is unavailable, or the thresholds are unreachable at acceptable cost. Narrowing is the most common outcome and the most valuable; stopping is a success that saved a build. Pilot agreements that follow are in how to structure an ai pilot agreement.
How do you run discovery in a few weeks?
- Week one: baseline extraction; correctness sessions; data access requests; integration inventory.
- Week two: first golden dataset labeled; candidate approaches run; sample data assessed.
- Week three: evaluation results by category; cost model; risk register; specification draft.
- Week four: specification finalized; decision meeting with evidence; staged funding plan.
Timeboxes vary with workflow complexity, but discovery that runs for months has become a pilot without a decision. Specification discipline is in the spec-driven development whitepaper.
What mistakes undermine discovery?
Estimating the baseline instead of measuring it; letting engineers define correctness without domain experts; testing feasibility with a demo on hand-picked examples; assuming data access; skipping integration inventory; and treating stop or narrow as failure. Each reappears as a stalled build. Prioritization upstream of discovery is in how to prioritize ai use cases.
How FISTA Solutions runs discovery
FISTA Solutions runs discovery as a fixed-price phase: baseline measurement, correctness sessions with client domain experts, a first golden dataset, feasibility evaluation across candidate approaches, data and integration verification, cost model, risk register, and a specification with a proceed, narrow, or stop recommendation, all delivered in weeks. The AI enablement practice leads discovery, forward deployed engineers carry the specification into delivery, and AI agents supplies the systems. The record behind the approach is 150+ projects for 50+ companies.
To know whether and how before you fund the build, message FISTA on WhatsApp, or read how to write an ai spec for the document discovery produces.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How long should AI discovery take?
Typically a few weeks for a bounded workflow: enough to measure the baseline, build a first golden dataset with domain experts, run candidate approaches against it, verify data and integration, and write the specification. Discovery that runs for months is a pilot without a decision.
02What does discovery deliver?
A measured baseline, a specification with scope, acceptance criteria by category, failure behavior, and autonomy level, a first golden dataset with evaluation results, a data and integration readiness report, a cost model at expected adoption, a risk register, and a recommendation to proceed, narrow, or stop.
03How is feasibility tested?
By running candidate models and approaches against the first golden dataset with calibrated grading, by category, and comparing results with the draft acceptance thresholds. Demos show possibility; evaluation shows whether the thresholds are reachable and at what cost.
04Who should be involved?
The process owner and domain experts who define correctness, an AI engineer who runs evaluation, a data owner who verifies access and permissions, an integration owner for the target systems, risk or compliance for tiering, and the sponsor who makes the decision.
05Is stopping a failure?
No. Discovery that finds the baseline too small, the data unavailable, or the thresholds unreachable has saved the cost of a failed build. Narrowing scope to what evaluation shows is reachable is the most common and most valuable outcome.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.