HomeCase Studies

AI & ML Solutions

Case Study · AI & ML Solutions

68% Fewer Support Tickets. One AI Agent.

How we built a production-ready, context-aware AI support agent for a US SaaS company — from zero to live in 14 weeks.

Service

AI & ML Solutions

Location

USA

Duration

14 weeks

Year

2025


Industry

SaaS / Customer Support

Project Overview

A fast-growing B2B SaaS company in the USA was drowning in support tickets. Their team of six agents was spending 24+ hours a week on repetitive L1 queries — password resets, billing questions, feature explanations — while real technical issues went unresolved for days. They approached Futurise Solutions to build an AI-powered solution that could deflect routine queries without sacrificing the quality of the customer experience.

Key Insight: Clean Data Over Big Models

We discovered that prompt engineering and clean chunking of documentation had a 5x larger impact on accuracy than using larger, more expensive LLM models. Structuring context properly eliminated 98% of hallucinations.

The Challenge

The Problem

The company's existing help centre had over 800 articles, but only 12% of users found answers before opening a ticket. Their previous chatbot attempt used simple keyword matching — it was brittle, frequently wrong, and trained users to bypass it entirely. The team needed an AI that could reason over their documentation, understand nuance, and give accurate answers while maintaining a strict guardrail against hallucination.

800+ documentation articles with no unified retrieval layer

Previous keyword chatbot had a 34% accuracy rate — users had abandoned it

Support team spending 24 hrs/week on L1 queries, delaying L2/L3 resolution

No visibility into which questions were being asked most

GDPR and SOC 2 compliance requirements for data handling

Our Solution

Our Approach

We designed and built a full RAG (Retrieval-Augmented Generation) pipeline backed by a purpose-built knowledge base. Rather than connecting an off-the-shelf chatbot, we engineered the system to understand the client's specific product context — integrating live ticket data, documentation, and CRM metadata into a unified vector store queried in real-time.

01

Knowledge Architecture

Chunked and embedded 800+ documentation pages, release notes, and SOPs into a MongoDB Atlas vector index using `text-embedding-004`. Built a metadata taxonomy for accurate retrieval.

02

RAG Pipeline

Semantic search retrieves the top-5 most relevant chunks per query. A Gemini 2.5 Flash model assembles grounded answers with strict context windows to prevent hallucination.

03

Guardrails & Escalation

Confidence scoring below a 0.75 threshold triggers a live-agent handoff. A custom intent classifier routes billing and account queries to the appropriate internal API.

04

Analytics Dashboard

Built a real-time dashboard showing deflection rate, query clusters, missed topics, and CSAT scores — giving the support team actionable data to improve the knowledge base weekly.

Execution Timeline

Our Step-by-Step Process

Weeks 1-3

Discovery & Data Prep

Auditing help center docs, cleaning metadata, and formatting data for retrieval models.

Weeks 4-7

RAG & LLM Prototyping

Setting up MongoDB Atlas Vector search, building the semantic search layer, and refining Gemini prompts.

Weeks 8-11

Guardrails & Integration

Creating confidence threshold routing, implementing fallback to live agents, and integrating Stripe API.

Weeks 12-14

Testing & Rollout

Shadow-testing with real support agents, tuning responses, and moving to 100% customer-facing traffic.

Tech Stack

  • Google Gemini 2.5 Flash
  • text-embedding-004
  • MongoDB Atlas Vector Search
  • Node.js / Express
  • React
  • Vercel
  • Intercom API
  • Stripe API (billing queries)

Results

Measurable impact — verified outcomes

68%

Reduction in support tickets

Within the first 90 days of deployment

64%

L1 query deflection rate

Queries answered without human involvement

24 hrs

Team hours saved per week

Redirected to complex technical escalations

4.5 / 5

AI response CSAT score

Measured via post-chat survey

Retrospective

Lessons Learned & Refinements

Dynamic Chunking Wins

Standard character-count chunking broke code blocks and step-by-step instructions. We switched to syntax-aware chunking to preserve procedural information.

Human-in-the-Loop Backup

Initially, the agent tried to handle too much. Creating an early handoff trigger for frustrated sentiment preserved customer CSAT scores.

"We expected the chatbot to handle maybe 30% of tickets. It handled 64% from week one, and the accuracy was genuinely impressive. The Futurise team understood our product deeply — this wasn't a generic AI deployment."

Marcus T.

VP of Customer Success, US SaaS Platform

Frequently Asked

Trending Questions & FAQs


Want similar results?

Let's build your success story.

Book a free 45-minute strategy call. We'll scope your project, recommend the best approach, and give you a clear roadmap — no strings attached.