AI economics and optimisation

Make AI cheaper, faster and commercially viable.

We reduce the cost of building and running AI — combining AI engineering, infrastructure optimisation and cost analysis to improve model usage, architecture and unit economics, without compromising quality.

Understand your AI spend. Eliminate waste. Improve unit economics.

Engineering delivered for

  • Sky
  • BBC
  • TfL
  • IPF

Most companies cannot say what they spend on AI. Fewer still know what each useful outcome should cost.

The spend is real but scattered: direct provider cards, assistant seats bundled inside Microsoft 365 or Google Workspace, Azure OpenAI and Bedrock buried in the cloud bill, and tools bought on personal expenses. Until it is one number, attributed to teams and use cases and measured against the quality it delivers, there is no way to separate necessary cost from waste.

More than cheaper tokens

Switching supplier is the easy part.

Cheaper tokens are a small share of the saving. The rest comes from changing how the system works — what runs, on which model, how often, and whether it needed to run at all. That is engineering, and it is most of what we do.

  • Inference and infrastructure cost
  • Cost per successful outcome
  • Model and provider selection
  • AI project ROI
  • Scaling economics
  • Wasteful pilots and unused licences
  • Performance, latency and reliability

The problem

Where the money goes.

Organisations are adopting AI faster than they are costing it. The same patterns appear again and again — and almost all of them are fixable without reducing what the system can do.

Typical today

  1. Spend scattered across cards, cloud bills, seats and expense claims
  2. Seats bought every month that nobody has opened
  3. Expensive models doing work smaller models could handle
  4. Every request carrying unnecessarily large context
  5. Agentic workflows making repeated or uncontrolled calls
  6. No caching, batching or intelligent model routing
  7. Cloud GPUs and inference infrastructure underutilised
  8. Self-hosting adopted without costing the operational overhead
  9. LLM calls doing work deterministic code could do
  10. Success measured in tokens, not cost per outcome
  11. No budgets, alerts, ownership or financial KPIs

After optimisation

  1. One number for AI spend, attributable by team, tool and use case
  2. Licences matched to the people who actually use them
  3. Each task routed to the cheapest model that passes evaluation
  4. Context trimmed to what measurably improves the answer
  5. Agent loops bounded, deduplicated and observable
  6. Caching, batching and routing built into the request path
  7. Utilisation matched to real demand curves
  8. Hosting decided on total cost, not on principle
  9. Deterministic work moved out of the model
  10. Cost per successful outcome tracked per use case
  11. Budgets, alerts and named ownership in place

The engagement

The AI Cost Audit.

A fixed-price engagement, typically two to four weeks, covering nine workstreams. It ends with a costed optimisation plan — not a set of observations.

  1. Spend and usage analysis

    Pull together provider billing, cloud cost exports, licence reports and expense claims into one picture: expenditure by supplier, model, product, team and use case, so every pound is attributable to something the business recognises.

  2. Seat and licence reconciliation

    Compare seats paid for against seats actually used — Copilot, Gemini, Cursor, Perplexity and the rest — from the admin reports you already have. Unused seats are usually the fastest money back on the table.

  3. Shadow AI and unmanaged spend

    Find the AI tools bought outside procurement — on personal expenses, team cards and free tiers that quietly became paid. You cannot govern or cost what nobody has counted.

  4. Model assessment

    Test whether work can move to cheaper models without unacceptable quality loss — examining routing, fallbacks, smaller models, open-source models and specialist models.

  5. Prompt and context optimisation

    Reduce unnecessary instructions, repeated context, excessive retrieval, output length and redundant agent calls — the quiet majority of most token bills.

  6. Infrastructure review

    Compare API, managed cloud and self-hosted deployment — including GPU utilisation, scaling, quantisation, concurrency and the operational overhead each option carries.

  7. Architecture review

    Identify opportunities for caching, batching, asynchronous processing, pre-computation, deterministic code and removing LLM calls that were never needed.

  8. Fine-tuning and retrieval assessment

    Determine whether fine-tuning, distillation or better retrieval would genuinely reduce recurring inference cost — or whether it would simply add complexity.

  9. Commercial and governance review

    Review contracts, licences, usage controls, budgets, monitoring and accountability, so the savings hold after we leave.

The deliverable

A financial and technical optimisation plan.

Every recommendation is quantified, ranked and tied to an implementation effort and a quality risk — so you can decide what to do on commercial terms, not technical ones.

Illustrative optimisation plan: recommendations with current cost, expected cost, saving, effort and quality risk
Recommendation Current cost Expected cost Saving Effort Quality risk
Reclaim assistant seats with no recent activity £X£Y £Z Low Low
Consolidate duplicate and overlapping AI subscriptions £X£Y £Z Low Medium
Route simple requests to a smaller model £X£Y £Z Low Low
Reduce average context size £X£Y £Z Medium Low
Cache repeated outputs £X£Y £Z Low Low
Bound and deduplicate agent call loops £X£Y £Z Medium Low
Move batch workload to self-hosting £X£Y £Z High Medium

Illustrative structure. Your plan carries your own figures, measured from your own usage data.

The aim is not fewer tokens. It is a lower cost per accurate, useful business outcome — and every recommendation is verified by evaluation before we propose it.

Anything that saves money by quietly degrading quality is not a saving; it is a deferred cost. Every change we recommend is tested against an evaluation set built from your own workload, so you can see exactly what the cheaper configuration does and does not do.

Independence

We have nothing to sell you but the saving.

We are not trying to sell you a particular model, cloud or platform, and we take no vendor commission. We determine which combination delivers the required quality at the lowest sustainable cost.

That means being honest in both directions. Self-hosting, fine-tuning and smaller models are not automatically cheaper — once utilisation, engineering time, monitoring and GPU operations are included, the major APIs are often substantially cheaper than running your own.

Sometimes the right answer is to consolidate onto one supplier. Sometimes it is to leave the architecture alone and fix the commercials. We will tell you when there is less to save than you hoped.

The recommendation follows the numbers.

Cheaper, faster, and provably as good.

  • Attributedspend mapped to teams, products and use cases
  • Verifiedevery change tested against evaluations
  • Independentno vendor, model or platform incentive
  • Durablebudgets, alerts and ownership that outlast us

How to work with us

Find it. Fix it. Keep it fixed.

Three engagements. Start with the audit — if the savings do not justify the next stage, we will say so.

Stage one

AI Cost Audit

Find and quantify the savings. Spend analysis, model and architecture assessment, infrastructure and commercial review, delivered as a ranked, costed optimisation plan.

Fixed price. Two to four weeks. You keep the plan whether or not we implement it.

Stage two

AI Optimisation Sprint

Implement the highest-value changes. Routing, caching, batching, context reduction, architecture changes and infrastructure moves — each shipped behind evaluations that prove quality held.

Scoped from the plan. Prioritised by saving against effort and risk.

Stage three

Continuous AI Cost Management

Monitor spend, quality and cost per outcome every month. New models, new usage patterns and new teams all move the economics — someone needs to be watching.

Monthly retainer. Reporting your finance team can read and your engineers can act on.

Who this is for

If nobody can name the number, start here.

  • Nobody in the business can give one figure for what AI costs
  • Spend is split across provider cards, cloud bills, seat licences and expenses
  • You pay for assistant seats you suspect are barely used
  • Teams are buying AI tools independently, outside procurement
  • You run AI in production and inference cost rises with customer growth
  • You built a prototype whose economics do not work at scale

You don't have to build AI to be overspending on it.

Buying it is enough. Most of what the audit finds early is bought, not built — seats nobody opens, overlapping tools, subscriptions renewing on a card no one owns. If you do run your own AI systems, the engineering workstreams go after inference cost as well, and that is usually where the larger savings sit. See the engineering capability.

Why our recommendations hold

We optimise AI systems because we build them.

Cost advice is only worth taking from people who could implement it. The same engineering depth that lets us ship agentic systems into production is what lets us find the waste inside them — and then remove it without breaking anything.

  • Agentic AI systems

    Agents that plan, call tools and recover from failure — and, just as importantly, that stop. Bounded loops, deduplicated calls and human hand-back are where runaway agent bills are won or lost.

    • Tool use
    • Loop control
    • Multi-agent
    • Human-in-the-loop
  • LLM engineering

    Model selection, routing, retrieval design, fine-tuning and structured output — built against evaluation suites, so a cheaper configuration can be proven equivalent rather than hoped equivalent.

    • Model routing
    • RAG
    • Evals
    • Distillation
  • AI infrastructure

    Inference, vector and data layers sized to the workload — self-hosted, cloud or hybrid. GPU utilisation, quantisation, concurrency and the honest total cost of running your own.

    • Inference serving
    • GPU utilisation
    • Quantisation
    • Hybrid hosting
  • Security & governance

    Threat modelling, least-privilege tool access, PII handling and audit trails — plus the budgets, quotas and usage controls that stop a single misconfigured job becoming a month's spend.

    • Access control
    • Budgets & quotas
    • Data residency
    • Auditability
  • Scale & performance

    Systems built for real traffic: national-scale user bases, broadcast peaks and bursty workloads. Load modelling, caching, batching and the observability that turns spend into a metric you can manage.

    • Load modelling
    • Caching
    • Batching
    • Observability
  • Deployment & operations

    CI/CD, infrastructure as code, staged rollout and monitoring for systems whose behaviour is probabilistic — so an optimisation can be shipped, measured and rolled back if the numbers disagree.

    • IaC
    • CI/CD
    • Staged rollout
    • Cost monitoring

Sectors

Engineers who already know your constraints.

Our team has delivered production systems inside regulated, high-traffic and safety-critical environments — so cost recommendations arrive already tested against the rules you actually operate under.

Finance

Document intelligence, risk and decisioning workflows and customer operations — where a cheaper model still has to survive audit and model governance review.

Transportation

Operational data at network scale: real-time information, asset and disruption workflows, and public-facing services whose demand curve is anything but flat.

Entertainment

Media and content platforms: metadata enrichment, archive search and personalisation, where volume makes a fraction of a penny per call a budget line.

Digital healthcare

Clinical and patient-facing systems where data protection, information governance and clinical safety constrain which economies are available at all.

Why Manchester AI

AI expertise. Applied to the invoice.

There is no shortage of presentations explaining what AI might do. Rather fewer people can tell you what yours costs, why, and what it should cost instead. Manchester AI combines architecture, security engineering and practical AI expertise with a straightforwardly commercial question: what is each useful outcome costing you?

Our team has delivered systems for organisations including Sky, the BBC, TfL and IPF — across finance, transportation, entertainment and digital healthcare, in environments where downtime, data handling and scale are not abstract concerns.

We still design and build AI systems end to end. Increasingly, the more valuable work is making the ones you already run cost a fraction of what they do today.

Start with the spend

What is your AI actually costing you?

Bring us a bill you cannot fully explain, a margin that worsens as you grow, or a prototype whose economics do not survive scale. The audit tells you what is avoidable and what it would take to remove.

Fixed price. Two to four weeks. You keep the plan.