· By Ricardo Torres Oliva
What a Director needs to know about AI Implementation
What a board member has to be able to ask before approving an artificial-intelligence initiative, and the answers that disqualify it.
Abstract
Approving an investment in artificial intelligence without being able to interrogate it turns the board into a signatory of what others decide, and since 2026 that signature carries fiduciary consequences. Written for directors and committees who approve AI budgets without executing them, the text brings together three instruments for immediate use: seven requirements every proposal must meet before it is signed, a thirteen-question due-diligence checklist for the vendor with the answer that disqualifies in each case, and the five-indicator dashboard that shows, months later, whether what was implemented actually works. It closes with what is not worth buying and with seven actions that can be executed at the next board session. None of them requires technical training; all of them can be verified by a third party.
A board does not implement artificial intelligence: it approves, demands and verifies. This document delivers the three tools of that craft — what to ask a vendor, what to require in writing before signing, and what evidence to use afterwards to confirm that what was implemented works. It does not teach technology. It teaches how to decide about it.
1. The problem
80% of C-suite leaders declare high confidence in their artificial intelligence strategy. 35% pass a foundational AI literacy assessment [source: The Phoenix Doctrine v1.1]. The distance between those two figures is the whole problem, and the doctrine has a name for it: executive cognitive dissonance.
That distance is paid for in money. 95% of generative AI pilots deliver no measurable return, and the documented cause is not technological immaturity: it is that the organization implementing them lacks the executive fluency to translate business objectives into viable architectures [source: The Phoenix Doctrine v1.1]. At market scale, that return gap is estimated at around 600 billion dollars [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era].
Until 2024 this was an efficiency problem. Since 2026 it is a fiduciary one. The legal doctrine that corporate governance literature calls the Expertise Trap holds that a director who does not understand the fundamental logic of the algorithmic tools shaping their decisions can be held personally liable for lack of diligence. Its counterpart is the Business Judgment Rule 2.0: legal protection is preserved only for those who can demonstrate critical interrogation of the tool, not acceptance of its output [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era].
The position of this document is simple. Buying artificial intelligence without the capacity to interrogate what was bought is not delegation. It is abdication with an invoice attached.
2. Why the usual approaches fail
Hiring technical talent and delegating. 38% of companies have appointed a Chief AI Officer, and their effectiveness is systematically undermined by fragmented reporting lines [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era]. The real shortage was never one of engineers. It is one of executives who know what to ask of them, how to evaluate them and how to integrate what they produce into business decisions [source: The Phoenix Doctrine v1.1]. An organization with good engineers and poor AI leadership produces brilliant pilots that never scale.
Buying a three-year transformation. This is the fragile middle: massive budget, long horizon, linear return promised. In a field that changes dramatically every six months, that project arrives obsolete halfway down the road [source: The Phoenix Doctrine v1.1]. The doctrine explicitly names as its enemy the consulting model that sells time and man-hours as the measure of value.
Waiting for clarity. Paralysis dressed up as prudence. The director who says "first I want to understand this properly" is usually rationalizing the fear of being wrong, and in this transition the error of not acting is greater than the error of acting badly and correcting course [source: The Phoenix Doctrine v1.1].
Confusing IT modernization with a strategic pivot. This is the most common literacy gap at board level: evaluating AI investment as a line of expenditure rather than as a change in the organization's set of capabilities [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era]. A migration is approved and recorded in the minutes as a transformation.
3. What to require in a proposal, before signing
Seven requirements. None of them calls for technical knowledge to be formulated, and all of them can be verified by a third party.
A baseline measured before anything is installed. The figure for the process as it stands today: cost, cycle time, volume, error rate, with method and date. Without a prior baseline, no later improvement can be demonstrated, and the conversation about return becomes a negotiation of anecdotes.
A success criterion written as a rate, not as an adjective. An agentic system is not deterministic: the same task can produce different results, which is why it is verified against a curated set of real cases with expected outcomes, one that yields a rate rather than a verdict [source: Evals Ingeniería de Confiabilidad y Evaluación de Sistemas Agénticos]. The operating definition worth adopting in the boardroom is hard and useful: a pilot is a system without evals — it worked in the demo because nobody measured the rest of the distribution.
Total cost, including what sits below the waterline. There are three items that proposals omit with regularity: retraining for model drift, in the order of 22% of ongoing overrun; the data and governance multiplier, which demands close to four dollars of investment in data foundations for every dollar put into models; and the drag of technical debt from integration with legacy systems, around 29% of the return [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era]. A price that only covers licenses and implementation is an incomplete price, not a low one.
Contractual optionality. Exit clauses, data portability, modular architecture, the capacity to change vendor in weeks rather than years. The temptation to commit to a single stack in exchange for a discount or "deep integration" is the most expensive trap of this era [source: DAL OS v1 0]. The structural freedom to choose when the moment to choose arrives is worth more than any discount for exclusivity.
Vendor skin in the game. A share of the fee tied to the result. A vendor who charges the same whether it works or not has already put a price on their own indifference. The doctrine states it the other way round, and more demandingly: one cannot demand skin in the game from the client and collect one hundred percent of the fee regardless of the result [source: DAL OS v1 0].
An immutable, readable log. Every execution must leave a trace: what context was loaded, what it decided, which tool it invoked with which arguments, what it cost, where a human intervened. That log has three uses — debugging the failure, learning from it, and defending oneself before an auditor or a regulator. A system without an immutable log is not auditable, and a system that is not auditable should not operate without direct human supervision [source: Evals Ingeniería de Confiabilidad y Evaluación de Sistemas Agénticos].
Real human oversight, not decorative oversight. A stop mechanism, the ability to overrule the machine's decision, and an approver with training, context and allocated hours. Decorative oversight — approvers without training or context who sign everything — is worse than no oversight at all, because it produces false evidence of control, and it is the first finding an auditor looks for [source: Organización Humano-Agente La Empresa Híbrida y el Rediseño del Trabajo].
4. Due-diligence questionnaire for an AI vendor
Thirteen questions. The third column is the useful part: it does not describe a bad answer, it describes the answer that ends the conversation.
| Question for the vendor | What it tests | Disqualifying answer |
|---|---|---|
| What is the baseline for this process, measured before anything is installed, and by what method? | Whether a measurement apparatus exists or only a narrative | "We measure it after deployment" · an estimate with no method and no date |
| What success rate does the system sustain on our own real cases, and on how many cases? | Whether there are evals or only a demo | A demonstration instead of a figure · "it works very well" · laboratory cases |
| Who judges whether it worked, and are they independent of whoever built it? | Structural separation between judge and builder | The same team that builds evaluates its own work |
| What does a successful task cost, counting the failed attempts? | Real economics versus brochure economics | Cost per license or per user as the only figure |
| What part of the price covers retraining for drift, data governance and integration with legacy systems? | Whether the price is complete | "Everything is included", with no breakdown of those three items |
| If we change vendor in nine months, what does the company take with it, in what format and in how much time? | Real optionality | "Your data is exportable", with no timeframe, no format and without the accumulated evaluation cases |
| What part of your fee depends on this working? | Skin in the game | None · a symbolic retroactive discount |
| Which of our processes ceases to exist when this goes into operation? | Whether it is a redesign or a layer on top | "None, this adds to what you already do" |
| Who in our company remains responsible for the system, and what must they know in order to be so? | Whether the capability is installed or rented | "We operate it for you", presented as an advantage |
| What record remains of each decision the system makes, and who on our side can read it without your help? | Auditability and forensics | A log accessible only from their platform · "the model is a black box" |
| How is it stopped, who stops it and how quickly? | Effective oversight | No defined mechanism · stopping depends on a support ticket |
| How many human interventions per day does it consume in steady state, and of how many minutes each? | Hidden cost of oversight | "It is fully autonomous" · they have not measured it |
| What went wrong in your last implementation and what did you change afterwards? | Operational honesty and learning | "Nothing has ever gone wrong for us" |
Rule of use: the first three questions are eliminatory. Without a baseline, a rate on the company's own cases and an independent judge, the remaining ten are decoration on top of a pilot.
5. How to know, afterwards, whether what was implemented works
"It works" is not a state. It is a rate that decays: 91% of machine learning systems experience measurable degradation in their first twelve months [source: The Phoenix Doctrine v1.1]. A single verification at project close measures the system's best day, not its behavior.
The minimum dashboard a board should request as naturally as it requests the financial report consists of five indicators [source: Evals Ingeniería de Confiabilidad y Evaluación de Sistemas Agénticos]:
Success rate per task on real cases, not on the original acceptance set. This is the number that governs how much autonomy the system deserves.
Human intervention rate. How much oversight it really consumes. This is the cost that never appears in the proposal.
Cost per successful task. Failures are paid for too. This figure, and not the platform's monthly invoice, is the one compared against the baseline.
Time to detection. How long the organization takes to notice a deviation. It defines the real operational risk.
Weekly drift. The trend of the four above. It is the early alarm.
To those five indicators it is worth adding four closing statements, which are harder to dress up than any metric [source: DAL OS v1 0]: real destruction was executed, with date and signature — a process ended, a contract was cancelled, a product was withdrawn; the system was left more antifragile, in terms the team itself can articulate without the vendor's help; people's capability rose in ways verifiable in practice; and the organization can operate without the vendor. An engagement that ends with the client saying "we need you for everything else" sold well and failed.
6. What not to buy
Do not buy IT modernization labeled as a strategic pivot. Do not buy three-year transformations with massive budgets and promised linear returns. Do not buy stack exclusivity in exchange for a discount. Do not buy one more pilot: if the previous one has no measured rate, the next one will not have one either. Do not buy one-off training events without reinforcement or evidence — the Research_Base names them, without euphemism, a skills placebo [source: Organización Humano-Agente La Empresa Híbrida y el Rediseño del Trabajo]. And do not buy autonomy over processes whose log cannot be read: delegation without a trace is not efficiency, it is exposure.
There is one purchase that does deserve to be defended in the boardroom and is almost never proposed: redesign, training and oversight. A portfolio that allocates 90% to licenses and models and 10% to people and processes is inverted relative to where the value sits, and it predicts the failure of the pilot [source: Organización Humano-Agente La Empresa Híbrida y el Rediseño del Trabajo]. The reference heuristic distributes success at roughly 10% algorithm, 20% technology and data, 70% people and processes; the source itself asks that it be read as an order of magnitude, not as a measurement.
7. What to do on Monday
- Request the inventory of AI systems in operation, with one name per system. Not a department: a person. An agent without an owner is shadow AI with a budget [source: Organización Humano-Agente La Empresa Híbrida y el Rediseño del Trabajo].
- Request the five indicators from section 5 for each one. If they do not exist, that is the board's first decision, and it is cheaper than any new project.
- Apply the questionnaire in section 4 to the proposal on the table today, not to the next one. Record in the minutes which questions went unanswered.
- Request the breakdown of the AI portfolio between licenses and models on one side, and redesign, training and oversight on the other. The ratio is the diagnosis.
- Demand the measured baseline of a single process before the end of the quarter. One process, chosen for its volume and its cost, not for its visibility.
- Put on record the basis of every AI-assisted recommendation the board receives. Traceability of the algorithmic recommendation is today a documented fiduciary obligation, not good practice [source: AI Literacy A Multi-Dimensional Analysis of Governance, Revenue Systems, and Epistemic Rigor in the Agentic Era].
- Execute one act of destruction, with date and signature. A process that ends, a contract that is cancelled, a report that is no longer produced. The difference between an organization that says it destroys and one that destroys is always that signature [source: DAL OS v1 0].
What we have not covered here
This document stops at the boardroom door. It does not cover what happens when the decision goes down the line: a decision approved here reaches senior management as a budget and a deadline, without the criteria by which it was taken, and each function completes it with whatever criterion it has to hand. That is a different problem, with its own document.
Nor does it cover the internal architecture of what is bought — agent anatomy, agentic security, context engineering, autonomy levels — which is the business of those who build, not of those who approve. It does not cover the regulatory regime applicable outside the European Union, whose AI Act is the reference used here because it is today the de facto global standard. And it does not cover how a board installs the capacity to sustain these questions without external assistance, which is the work of Phoenix PEEx, from VoltAi Academy.