Data Governance

AI Budget Planning for FY2027: A CFO's Playbook

FY2027 is the first fiscal year in which most CFOs will treat AI as a line item to be managed rather than a bet to be defended — and this playbook is about doing that without strangling the programs that actually work.

Key Statistics: IDC (2024) forecasts worldwide AI spending to reach approximately USD 632 billion by 2028. Gartner (2025) estimates that more than 40% of agentic AI projects will be canceled by the end of 2027 due to escalating costs and unclear business value. A widely cited MIT study (2025) estimated that about 95% of enterprise generative AI pilots produced no measurable P&L impact. McKinsey (2025) reports that while roughly 78% of organizations use AI in at least one function, only about a quarter attribute enterprise-level EBIT impact to it. Gartner (2024) predicted around 30% of generative AI projects would be abandoned after proof of concept by end of 2025. The planning conclusion is consistent across sources: the problem is not AI spend, it is unmanaged AI spend.

Why FY2027 AI Budgeting Is a Different Exercise

For two budget cycles, AI spending lived in a protected category. Boards understood it as strategic exploration; CFOs signed off on pilots the way they sign off on R&D — knowing most attempts fail, but that the option value justifies the cost. That era is ending, and the numbers explain why. Gartner (2025) estimates over 40% of agentic AI projects will be canceled by 2027; a widely cited MIT study (2025) put the share of generative AI pilots with no measurable P&L impact near 95%. When failure rates like that become public knowledge, "trust the exploration" stops being a viable budget posture.

FY2027 budgeting is therefore different in kind, not just degree. Three structural changes define it:

  • AI moves from project funding to run-rate visibility. What used to be a one-time pilot budget now includes recurring model, infrastructure, and platform costs that behave like software subscriptions — except they scale with usage, not seats.
  • Finance gets a seat at the architecture table. Decisions such as which models to route to, whether to self-host, and how to meter internal consumption now directly determine cost curves. These are finance questions as much as technology questions.
  • The unit of analysis shifts from "AI initiative" to "decision improved." The FY2027 test for funding is no longer "is this innovative?" but "does this measurably make a specific decision faster, cheaper, or better — and can we prove it with a number we agreed on in advance?"

McKinsey (2025) found that only about a quarter of AI-adopting organizations report enterprise-level EBIT impact. The implication for budget planning is uncomfortable but clarifying: most of what your organization currently spends on AI is, by the evidence, not yet producing returns. The playbook below is about finding which part is, funding it properly, and stopping the rest.

The Structural Split: Platform Versus Pilots

The first architectural decision of an FY2027 AI budget is not a technology decision at all — it is how much of the budget goes to experimentation versus how much goes to industrialization. Most FY2025 and FY2026 AI budgets were structurally skewed: 70–90% of spend sat in pilots, each with its own tooling, its own data pipeline, and its own evaluation standard. That structure produced exactly what it was designed to produce — many experiments and few production systems.

For FY2027, the healthier structure inverts the ratio for organizations past the discovery phase. A practical reference frame:

Budget componentTypical FY2026 shareRecommended FY2027 shareWhat it funds
Experiments and new pilots60–80%15–25%New use-case validation with fixed-price, time-boxed trials
Production platform and data readiness10–20%35–45%Semantic layers, permissions, connectors, integration (e.g., MCP-based), observability
Operations and governance5–10%20–30%Monitoring, evaluation harnesses, audit trails, model cost management
Skills and change adoption5–10%10–15%Training decision-makers and analysts on AI-delivered workflows

The logic behind the shift: platforms amortize across use cases, pilots do not. Every pilot that hard-codes its own data integration is a tax on every future pilot. The platform share of the budget is what makes the tenth use case cost 20% of what the first one did — and that cost curve, not any individual use case, is where durable value concentrates.

The failure mode to avoid is the mirror image: over-consolidating before any pilot has proven value. If your organization has not yet demonstrated a single AI use case with defensible unit economics, do not fund a platform at 40% — fund two or three falsifiable pilots and let the evidence set the allocation. The split above assumes discovery is done; if it is not, finish it cheaply first.

New Line Items: Modeling LLM and Token Operating Costs

The FY2027 budget contains cost lines that had no equivalent two years ago, and most organizations are still modeling them badly. Token-based model costs behave unlike any legacy IT cost: they scale with conversational volume, are highly sensitive to architecture choices, and can swing 10x based on decisions made before the first invoice arrives.

Break the model cost line into four components:

  • Inference spend per use case. Not one "AI cost" line — one line per production use case, metered. A customer-service summarization use case and an analyst copilot have wildly different token profiles and must be budgeted separately.
  • Model tiering. Route simple queries to cheaper models and reserve premium models for tasks that need them. In practice, well-architected deployments route 60–80% of requests to small or mid-tier models; the cost delta between routing everything to a frontier model and tiering intelligently is routinely 3–5x on the same workload.
  • Context and retrieval costs. Every query that drags a large context window through the model pays for it. Optimizing retrieval — sending less, more relevant data — is a cost-control lever, not just a quality lever.
  • Evaluation and guardrail costs. Running evals, red-teaming, and monitoring generates its own inference spend. It is small relative to production traffic but must be named, or it gets silently absorbed and resented.

Two budgeting rules keep this manageable. First, require every AI line item to have a usage cap and a budget alert at 80% — surprises, not absolute levels, are what damage finance's relationship with AI programs. Second, negotiate committed-use pricing once a use case's monthly volume is stable; frontier-model pricing has fallen sharply through 2024–2026, and organizations that renegotiated mid-cycle captured substantial savings. IDC's (2024) spending forecasts, which show platform and services spend growing faster than raw model APIs, reflect exactly this maturation.

Unit Economics: The CFO's Core AI Metric

Every FY2027 AI funding decision should reduce to one question: what does one unit of this workload cost, and what is that unit worth? The discipline is borrowed from SaaS and marketplace businesses, and it transfers cleanly.

Define the unit per use case. For a conversational BI deployment, the unit might be "one data question answered with a permissioned, sourced answer." For a document-processing pipeline, "one contract extracted and validated." For a service copilot, "one customer interaction assisted." Then build the simple table finance can actually govern:

Use caseUnitUnit cost (modeled)Baseline cost of the same unit, manualMonthly volumeMonthly net benefit at steady state
Conversational BI (IM-native)Answered data questionHKD 0.4–1.2HKD 15–30 (analyst time)40,000Meaningful, staff-time led
Contract extractionDocument processedHKD 2–5HKD 60–120 (reviewer time)3,000Meaningful, cycle-time led
Service interaction copilotAssisted interactionHKD 0.8–2.0HKD 8–15 (handle-time delta)25,000Meaningful, quality + handle-time led

The numbers above are illustrative modeling anchors, not benchmarks — but the shape is the point. Note two properties. First, unit costs are cents against baseline costs of dollars; the economics of AI use cases are almost always dominated by the human time they displace or accelerate, not by token prices. Second, volume drives everything: a use case with excellent unit economics at 40,000 monthly queries is irrelevant at 400, which is why adoption (not accuracy) is the most common reason AI programs miss their business case.

McKinsey's (2025) finding that only ~25% of AI adopters report enterprise-level EBIT impact is, in unit-economics terms, a volume problem more than a cost problem. The FY2027 budget should therefore fund adoption mechanisms — embedding AI where work already happens, for instance in messaging platforms — with the same seriousness as the technology itself.

Defunding Zombie Pilots: A Principled Cull

Every enterprise carrying more than five AI initiatives has zombies: pilots that finished their evaluation window months ago, produce no decision-relevant metric, and survive because nobody owns the decision to stop. Gartner (2024) predicted ~30% of generative AI projects would be abandoned after proof of concept; the FY2027 budget cycle is where your organization formally performs that abandonment, on purpose, rather than by attrition.

A principled cull has three steps. Step one: publish the criteria before reviewing anything. Suggested bars: each initiative must name its owner, the decision it improves, its monthly run-rate cost, and one pre-agreed measurable result. Anything missing two of the four enters the review queue. Step two: review against those criteria in one sitting. The review is administrative, not political — the criteria do the deciding. Step three: sunset with a deadline, and reallocate visibly. Freed budget should flow to the platform and operations lines described above, in the same cycle, so the cull visibly funds something rather than just removing something.

Two failure modes deserve names. The "sunk-cost zombie" survives because its sponsor has reported on it for a year; the fix is requiring a falsifiable result, not progress narratives. The "stealth zombie" survives because its run-rate is hidden inside a cloud bill or a team's operating budget; the fix is the complete inventory — every AI initiative, one page, four fields. A widely cited MIT study (2025) put the share of valueless generative AI pilots near 95%; your organization's number will be lower, but it will not be zero, and the FY2027 budget should be built as if it will not be small.

Procurement Checkpoints: Buying AI Without Buying Surprises

AI procurement has features that classic software procurement handles poorly: usage-based pricing that can scale unpredictably, model behavior that changes with vendor updates, and data flows that cross organizational boundaries. Five checkpoints, written into the standard process for FY2027, close most of the gaps:

  • Checkpoint 1 — Data flow disclosure. Where does data transit, where is it stored, is it used for vendor model training, and how is access revoked? Make the answer a contract annex, not a sales-deck assurance.
  • Checkpoint 2 — Permission architecture. The system must enforce the requesting user's data permissions at query time. Any tool that answers from a god-mode service account is disqualified for anything touching regulated data — IBM (2025) found 97% of organizations with AI-related breaches lacked AI access controls.
  • Checkpoint 3 — Priced usage model. Vendor pricing must be expressible as cost per unit of the workload (per query, per document), with volume bands and caps. If a vendor cannot state unit economics, the buyer will own the surprise.
  • Checkpoint 4 — Evaluation rights. Contractually reserve the right to run pre-agreed accuracy tests against your own data before renewal. Model updates can silently change behavior; the renewal decision needs evidence.
  • Checkpoint 5 — Exit and portability. Define what happens to conversation logs, embeddings, and configuration on termination. This is cheap to negotiate at signature and nearly impossible to renegotiate later.

For time-boxed purchases — pilots especially — the structure finance should prefer is small, dated, and falsifiable. A fixed-price two-week pilot in the HKD 25k range that tests a specific workflow against pre-agreed metrics is a better financial instrument than a six-month exploratory engagement at ten times the cost, because the downside is bounded and the information yield per dollar is higher.

Value Gates: Making AI Spend Falsifiable

The last piece of the playbook is the mechanism that makes everything above enforceable: pre-agreed value gates attached to each tranche of AI spend. A value gate is a commitment written before deployment — if the metric misses the bar at the checkpoint date, the next tranche does not release. It converts AI budgeting from a narrative exercise into a controlled experiment, which is the only form finance can govern at scale.

Design rules for value gates:

  • Metrics must be decision-level, not model-level. "95% answer accuracy" is a model metric; "store managers use it daily and reorder decisions happen 2 days earlier" is a decision metric. Fund the second.
  • Baselines are measured, not asserted. Measure the current cost and cycle time of the target decision in the two weeks before deployment. A gate with an unmeasured baseline is a gate that can be renegotiated — which means it is not a gate.
  • Checkpoints are short. Thirty to sixty days, aligned to how fast usage data accumulates. Annual value gates are where AI programs go to die quietly.
  • The consequence is real. A gate that misses must release nothing — not "with exceptions," not "pending a plan." One enforced gate buys more credibility for the whole AI portfolio than ten gates that bend.

McKinsey (2025) and IDC (2024) both point the same direction: AI value is concentrating in organizations that industrialize measurement. The value-gate mechanism is how a CFO imports that discipline without needing to personally evaluate retrieval architectures — the gates do the governing.

Who Owns What: Governing the Budget Process Itself

The mechanisms above — the platform/pilot split, unit economics, value gates — fail quietly when nobody owns the process that enforces them. FY2027 is the year to name that owner and write down the operating rhythm, because AI spending now crosses enough organizational boundaries that leaving it unassigned guarantees drift.

A workable ownership model keeps roles narrow:

  • The CFO owns the gates. Not the metrics themselves — business owners propose those — but the enforcement: whether a missed gate stops money. This is the one accountability that cannot be delegated downward without dissolving.
  • The CIO or CDO owns the platform line. Data readiness, connectors, permissions, and the delivery surface (including IM-native channels) are infrastructure decisions with 3–5 year cost implications; they should not be re-litigated per use case.
  • Business owners own unit economics. The head of supply chain, not IT, declares what a faster replenishment decision is worth. Finance validates the arithmetic; it does not invent the value.
  • Internal audit owns the governance checklist. Data flow disclosure, access controls, and audit trails get reviewed by the function whose finding actually hurts — which is what makes the checklist credible rather than ceremonial.

The operating rhythm is light but non-negotiable: a monthly one-hour AI spend review covering unit costs against budget, active usage against targets, and gate status per use case; a quarterly reallocation decision where freed budget moves between lines. McKinsey's (2025) finding that only ~25% of AI adopters report enterprise-level EBIT impact is, in many cases, an ownership failure before it is a technology failure — programs with no named owner drift toward exactly the unmeasured spend the 75% exhibit.

One caution: do not create a new committee for this. The monthly review should live inside the existing IT or transformation steering forum, with AI as a standing agenda item. New governance bodies multiply meetings without multiplying decisions, and FY2027 budgets fund decisions.

An Illustrative FY2027 Allocation

The table below sketches a reference allocation for a mid-size enterprise (roughly HKD 12–18M annual AI budget) that has completed discovery and has 2–3 use cases ready to industrialize. It is a starting point for negotiation, not a prescription.

Budget lineShareExample itemsGoverning value gate
Production platform & data readiness40%Semantic layer, permission-aware connectors, IM-native delivery (WeChat Work/DingTalk/Feishu/WhatsApp/Teams), observabilityTenth use case deploys at ≤30% of first use case's marginal cost
Industrialized use cases (run)25%2–3 production workloads: inference, integration, supportUnit cost per workload within ±20% of modeled budget, monthly
Governance, security, evaluation15%Audit trails, access controls, eval harnesses, monitoringZero critical audit findings; evals run monthly per use case
New experiments15%Fixed-price, time-boxed pilots (e.g., 2-week HKD 25k format)Each pilot passes or fails a pre-agreed 30-day metric
Skills & adoption5%Training, workflow redesign, champion networksActive-usage rate ≥60% of licensed users by day 60

Three notes on reading the table. First, the experiment line is capped by design — this is the structural protection against the pilot sprawl that consumed FY2025 and FY2026 budgets. Second, the adoption line is small in money but outsized in leverage; McKinsey's (2025) 25% figure is substantially an adoption failure, and five percent of budget spent on workflow redesign is often the highest-return line in the table. Third, every line has a gate — because in FY2027, the CFO's job is not to decide whether the organization believes in AI, but to ensure every remaining dollar of AI spend is still earning its place.

Frequently Asked Questions

For organizations past the discovery phase, a practical reference split is roughly 15–25% for new experiments, 35–45% for the production platform and data readiness, 20–30% for operations and governance, and 10–15% for skills and adoption. If no use case has yet demonstrated defensible unit economics, invert the emphasis: fund two or three falsifiable, fixed-price pilots first and let the evidence set the allocation.
The most commonly under-modeled items are token and inference operating costs that scale with usage, evaluation and monitoring spend, integration and permission work that each pilot would otherwise duplicate, and the internal staff time required to drive adoption. Budgeting per use case with usage caps and 80% budget alerts, and renegotiating committed-use pricing once volumes stabilize, keeps these lines predictable.
Publish the criteria before reviewing anything: every initiative must name its owner, the decision it improves, its monthly run-rate cost, and one pre-agreed measurable result. Review against those criteria in a single sitting, sunset failures with a deadline, and visibly reallocate the freed budget to platform and governance lines in the same cycle so the cull funds something rather than merely removing it.
Four properties: the metric is decision-level rather than model-level; the baseline was measured before deployment rather than asserted afterward; the checkpoint window is short, typically 30–60 days; and a miss actually stops the next tranche of funding. Gates that bend on first miss lose their governing power, while one enforced gate disciplines the entire AI portfolio.
Book a personalised demo

Ready to make your data auditable?

See how Beehive Strategy's conversational governance platform turns catalogues and lineage into answers your teams can query in plain language.

Book a Demo Explore the Solution
30%
Faster audit readiness
25%
Lower incident costs
40%
Less remediation time
2 wks
To a live catalogue