An AI product suite that saved £1.5M and lifted accuracy 30%
Shipped three production AI products, AI-Semantic Search, AI-Intake and AI-Summarizer, saving over £1.5M and lifting extraction accuracy by around 30%.
£1.5MCost savingsacross the AI suite
90%Faster searchAI-Semantic Search
+30%Extraction accuracyAI-Intake
3AI products shippedon AWS
What I owned
AI product strategy and roadmap for the suite
The HHH model-evaluation framework and accuracy targets
Data-training direction alongside engineering and data science
01
Context
The platform held large volumes of regulated safety documentation. Search was keyword-only, document review was manual, and a key-term intake step depended on extraction accuracy that had plateaued in the mid-60s.
02
Problem
Each of these was a direct drain on expert analyst time, the most expensive resource in the business, and a ceiling on how much the platform could scale per customer.
03
Discovery
Quantified the analyst time lost to slow search, manual summarisation and low-accuracy extraction.
Isolated three distinct AI surfaces, search, summarisation and intake, rather than one monolithic AI feature.
Defined success up front with an HHH evaluation framework, helpful, honest and harmless, with concrete accuracy targets.
04
Strategy
Ship three targeted AI products on AWS, each owning one bottleneck: AI-Semantic Search, AI-Summarizer and AI-Intake.
Match the architecture to the job: vector search for meaning, RAG for citable summaries, an NLP microservice for extraction.
Treat extraction accuracy as a first-class product metric, tracked release over release.
05
Execution
Built AI-Semantic Search on an embeddings and vector-search architecture, answering meaning-based queries instead of keyword matches.
Built AI-Summarizer on a RAG architecture, condensing 5,000 to 10,000 word documents to under 500 words with citable references, essential in a regulated domain.
Built AI-Intake as an NLP entity-extraction microservice, iterated across v1 and v2 with targeted data training.
Ran every model against the HHH framework before release.
06
Outcome
AI-Semantic Search reduced search time by 90% through meaning-based queries.
AI-Intake lifted extraction accuracy from the mid-60s to the upper-90s, around a 30% gain.
AI-Summarizer delivered cited summaries under 500 words, trusted enough for regulated review.
Together, the three products saved over £1.5M across the suite.
07
Key learnings
Define the evaluation framework before building the model, not after.
Accuracy is a product decision, not only an engineering one.
In regulated domains, citations are the feature that makes AI usable.
Next step
Looking for a PM who owns the strategy and ships the product?
Open to Senior PM, AI PM and Product Builder roles in London or UK-remote. Recognised Global Talent UK, with full working rights in the UK.