Case Study · AI & ML Solutions
68% Fewer Support Tickets. One AI Agent.
How we built a production-ready, context-aware AI support agent for a US SaaS company — from zero to live in 14 weeks.
Service
AI & ML Solutions
Location
USA
Duration
14 weeks
Year
2025
Industry
SaaS / Customer SupportProject Overview
A fast-growing B2B SaaS company in the USA was drowning in support tickets. Their team of six agents was spending 24+ hours a week on repetitive L1 queries — password resets, billing questions, feature explanations — while real technical issues went unresolved for days. They approached Futurise Solutions to build an AI-powered solution that could deflect routine queries without sacrificing the quality of the customer experience.
Key Insight: Clean Data Over Big Models
We discovered that prompt engineering and clean chunking of documentation had a 5x larger impact on accuracy than using larger, more expensive LLM models. Structuring context properly eliminated 98% of hallucinations.
The Challenge
The Problem
The company's existing help centre had over 800 articles, but only 12% of users found answers before opening a ticket. Their previous chatbot attempt used simple keyword matching — it was brittle, frequently wrong, and trained users to bypass it entirely. The team needed an AI that could reason over their documentation, understand nuance, and give accurate answers while maintaining a strict guardrail against hallucination.
800+ documentation articles with no unified retrieval layer
Previous keyword chatbot had a 34% accuracy rate — users had abandoned it
Support team spending 24 hrs/week on L1 queries, delaying L2/L3 resolution
No visibility into which questions were being asked most
GDPR and SOC 2 compliance requirements for data handling
Our Solution
Our Approach
We designed and built a full RAG (Retrieval-Augmented Generation) pipeline backed by a purpose-built knowledge base. Rather than connecting an off-the-shelf chatbot, we engineered the system to understand the client's specific product context — integrating live ticket data, documentation, and CRM metadata into a unified vector store queried in real-time.
Knowledge Architecture
Chunked and embedded 800+ documentation pages, release notes, and SOPs into a MongoDB Atlas vector index using `text-embedding-004`. Built a metadata taxonomy for accurate retrieval.
RAG Pipeline
Semantic search retrieves the top-5 most relevant chunks per query. A Gemini 2.5 Flash model assembles grounded answers with strict context windows to prevent hallucination.
Guardrails & Escalation
Confidence scoring below a 0.75 threshold triggers a live-agent handoff. A custom intent classifier routes billing and account queries to the appropriate internal API.
Analytics Dashboard
Built a real-time dashboard showing deflection rate, query clusters, missed topics, and CSAT scores — giving the support team actionable data to improve the knowledge base weekly.
Execution Timeline
Our Step-by-Step Process
Discovery & Data Prep
Auditing help center docs, cleaning metadata, and formatting data for retrieval models.
RAG & LLM Prototyping
Setting up MongoDB Atlas Vector search, building the semantic search layer, and refining Gemini prompts.
Guardrails & Integration
Creating confidence threshold routing, implementing fallback to live agents, and integrating Stripe API.
Testing & Rollout
Shadow-testing with real support agents, tuning responses, and moving to 100% customer-facing traffic.
Tech Stack
- Google Gemini 2.5 Flash
- text-embedding-004
- MongoDB Atlas Vector Search
- Node.js / Express
- React
- Vercel
- Intercom API
- Stripe API (billing queries)
Results
Measurable impact — verified outcomes
68%
Reduction in support tickets
Within the first 90 days of deployment
64%
L1 query deflection rate
Queries answered without human involvement
24 hrs
Team hours saved per week
Redirected to complex technical escalations
4.5 / 5
AI response CSAT score
Measured via post-chat survey
Retrospective
Lessons Learned & Refinements
Dynamic Chunking Wins
Standard character-count chunking broke code blocks and step-by-step instructions. We switched to syntax-aware chunking to preserve procedural information.
Human-in-the-Loop Backup
Initially, the agent tried to handle too much. Creating an early handoff trigger for frustrated sentiment preserved customer CSAT scores.
"We expected the chatbot to handle maybe 30% of tickets. It handled 64% from week one, and the accuracy was genuinely impressive. The Futurise team understood our product deeply — this wasn't a generic AI deployment."
Marcus T.
VP of Customer Success, US SaaS Platform
Frequently Asked
Trending Questions & FAQs
Want similar results?
Let's build your success story.
Book a free 45-minute strategy call. We'll scope your project, recommend the best approach, and give you a clear roadmap — no strings attached.
