The starting point nobody wants to admit
In 2026, most boards in Spain have the word AI on their agenda. Few have a question more concrete than "what do we do with this?". The most common result: an external consultancy that delivers 80 slides, an internal project that dies in the second sprint, or a ChatGPT pilot in the marketing department nobody can tell is working.
This guide gets down to the ground. Five concrete steps, with timings and estimated cost, to move from "we have to do something with AI" to "process X saves Y hours and costs Z euros a month".
Step 1 · Map processes before choosing models
The most expensive mistake we have seen in AI projects is starting with the tool. Someone decides that Claude or GPT-4 are "the future" and kicks off a pilot without first mapping which concrete process they want to attack.
The right order:
- List repetitive, high-volume tasks. Not strategic projects — tasks. "Categorise 200 tickets a day", "produce 50 weekly reports", "answer 30 standard emails a day".
- Estimate the current cost. Hours × cost/hour × frequency. If a task costs less than €500/month in human hours, it probably is not worth automating yet.
- Measure rules and variance. A task with clear rules and little variance between cases is an ideal candidate. A task with many edge cases and critical human judgement is better left for later.
The output of this step is a prioritised list with 5–10 candidates, each with its potential saving and its estimated complexity. Without that list, everything else is noise.
Step 2 · Choose the right pilot (not the sexiest)
Once you have the prioritised list, the temptation is to attack the most visible or most impressive process. Wrong. For the first pilot, choose the case that has:
- Medium-to-high ROI (it does not have to be the highest)
- Low risk (if the agent gets it wrong, the impact is limited)
- Fast feedback (you can tell whether it works in 4–6 weeks, not 6 months)
- Internal support (someone on the team wants it to go well and can give it time)
In practice, the pilots that work tend to be email categorisation, extracting data from documents, generating drafts of standard replies, or classifying level-1 support tickets. Little glamour, high impact.
Step 3 · Minimum viable architecture
An AI agent in production needs five minimum components:
- Model (Claude, GPT, Llama). A technical decision, not a marketing one.
- Authorised tools (internal APIs, databases, email).
- Guardrails: what it can and cannot do. By technical contract.
- Auditable log: every decision gets recorded.
- Kill switch: a button that stops the flow in seconds.
Any agent in production without these five components is an incident waiting to happen. It is not negotiable.
Step 4 · Measurement before and after
Define the metrics before starting the pilot. Afterwards is too late: confirmation bias guaranteed.
Typical metrics:
- Hours/week saved (measured against a human baseline)
- Error or rejection rate (the agent proposes something — does it get approved?)
- Cycle time (from the task arriving to it being closed)
- Marginal cost (model + infrastructure cost per operation)
- Human intervention (% of cases that required a manual adjustment)
A pilot with no metrics is an anecdote. With metrics, it is a replicable business case.
Step 5 · Scale without breaking
If the pilot works, the instinct is to extend it in all directions at once. Bad plan. Scaling has three dimensions and they have to be attacked one at a time:
- Volume: going from 100 cases/day to 1,000. It requires optimising the model cost and monitoring.
- Variety: adding edge cases to the scope. It requires more guardrails and testing.
- Geography / team: extending to other departments. It requires training and adjusting to their specific processes.
Scaling all three at once guarantees something breaks and trust in the system goes with it. One at a time, with a decision gate between each, works.
What it really costs
A realistic range for a mid-sized Spanish company (50–500 employees) implementing its first AI use case in 2026:
- Initial diagnosis: €6,000 – 12,000
- Working pilot (4–6 weeks): €15,000 – 35,000
- Scaling to production: €30,000 – 80,000
- Monthly operation (model + infrastructure + support): €1,500 – 8,000/month
With a well-chosen pilot, the typical payback is 3–6 months. If someone promises you a return in under 2 months, be suspicious.
The 5 mistakes we see repeated
- Starting with the tool. "Let us set up an internal ChatGPT" before knowing which problem it solves.
- Confusing a POC with a pilot. A script that works in a demo is not a system in production.
- Automating broken processes. If the process is badly defined, automating it only amplifies the problem.
- No compliance from day 1. Adding GDPR and the AI Act at the end is 10× more expensive than building them in from the design.
- No agreed metrics. With no clear baseline, there is no way to know whether what you built contributes anything.
The honest shortcut
If you want to avoid 80% of the typical mistakes, an external diagnosis of 2–3 weeks gives you a prioritised map of opportunities, identifies the candidate processes and gets you out of blank-page paralysis. It costs between €6,000 and €12,000. It saves months of going in circles.
We do that diagnosis. Details here. Hiring us is not essential, but hiring someone who does not only sell the tool — someone who comes with a method — is.