Frequently asked

Questions,
answered.

The things people ask before they bring local AI in-house — hardware, models, data, compliance and how we work. Can't find yours? Ask us directly.

The basics
What does "local AI" actually mean?

AI that runs on hardware you control — or in UK datacentres you can audit. The model, your data, and the inference all stay inside a boundary you define. No per-token billing from a foreign cloud, no documents shipped abroad for processing, and the system keeps working when the internet doesn't.

Do we need to buy hardware first?

No. There are three paths: run it on hardware you already have, buy a right-sized box (we'll tell you exactly what, at what price, before you commit), or rent from our Mac hosting line-up — a fleet of Apple silicon Macs and Mac Studios in UK datacentres, which make excellent local-LLM machines.

Do you only work with Apple hardware?

No. Apple silicon is excellent for many local-LLM workloads — and we host our own fleet of them — but it's not the only answer. The right hardware depends on the use case: some projects are better suited to Linux servers, NVIDIA GPUs, or high-memory workstations. We size hardware around the models and workloads, on a project-by-project basis, and we'll tell you honestly if a Mac isn't the right tool for the job.

Who are you, and where are you based?

We're MacHive — a UK-based team: we host fleets of Apple silicon Macs and Mac Studios in UK datacentres (see machive.co.uk), and we consult on local AI on top of that fleet. Remote-first, working across the UK, on-site when it helps. If you're based in the UK, your data can stay in the UK.

How big is a typical engagement?

Most people start with a free phone call, then a two-week Discovery (£1,800, fixed scope) that ends with a firm quote for the Build. A typical private-LLM + RAG build runs four to eight weeks. Larger programs get phased — always with a written scope and a price before we start.

Models & technology
Which models do you work with?

Open-weights families — Llama, Qwen, Mistral, Gemma and the like — quantised and served with the right tooling for the hardware. We pick the model to fit your quality and latency budget, not the other way round.

Why not just call a cloud API?

Honest answer: for many things, that's fine — and we'll tell you if it is. Local starts to make sense when the data is sensitive, when volumes make per-token billing painful, when latency matters, or when you need the system to keep working offline.

Will it work as well as the big hosted models?

For narrow, well-scoped tasks — document search, drafting in your own domain, classification, support triage — a well-tuned local model plus good retrieval closes most of the gap. We benchmark on your data during Discovery, so you can decide with evidence rather than vibes.

What happens when better models come out?

Open weights mean there's a new strong model roughly every few months. On a retainer we run the refresh for you: re-evaluate, canary it against your data, and switch when it wins.

Data & compliance
Does our data leave the building?

No — that's the point. Documents, embeddings, prompts and completions stay inside your network or your UK racks. We document exactly what crosses any boundary, and by default it's nothing.

What about GDPR?

Local inference removes most of the third-party-processor questions: no foreign cloud sees your prompts or documents. We put the architecture and data flows in writing so they slot straight into your DPIA.

Can we keep using our existing apps?

Yes — the stack fronts an OpenAI-compatible API, so most existing integrations work with a change of base URL and a new key.

Working with us
How long do projects take?

Discovery is two weeks. A typical build is four to eight weeks from kickoff to handover, depending on the document set and integrations.

What does it cost?

It starts with a free phone consultation. Discovery is £1,800 (credited in full to a build), and both the build and the managed retainer are priced on your requirements — typically £8,500+ for builds and £1,800–£2,400/month for retainers, fixed once quoted. Everything is in writing, ex VAT — the details are on the ways of working section.

What support do we get afterwards?

Yes — ongoing monthly support is available as a retainer, priced on what you need: monitoring, model refreshes, a monthly improvement sprint, new document and data onboarding, and priority support with a one-working-day response. Or we can hand over the keys and docs and you run it yourself.

Still wondering something?

If your question isn't here, just ask — we usually reply within one working day.

Talk to us