AI Restaurant Ordering
A waiter that answers menu and allergen questions from retrieval rather than model memory, on a locally hosted LLM.
Results
What it did
The problem
Why it needed building
Restaurant ordering systems are rigid — fixed menus, no dietary intelligence, no upselling. Customers want to ask natural questions like "What's good for someone with gluten allergies?" and get intelligent, context-aware responses. The system needed to run with low latency, handle concurrent conversations, and work with frequently changing menus.
The approach
How it works
Designed a full-stack system with vLLM for local LLM inference (eliminating API costs at scale) combined with a RAG pipeline backed by a vector database for menu retrieval. The conversation engine maintains multi-turn context, understands dietary restrictions, suggests complementary items, and handles payment handoff. WebSocket-based real-time streaming delivers responses as they generate.
- User Message
- RAG Retrieval
- Context Assembly
- LLM Inference
- Payment Handoff
- 01User MessageNatural language input via WebSocket connection
- 02RAG RetrievalVector DB queries menu items, allergens, and pairings
- 03Context AssemblyConversation history + menu data + dietary preferences
- 04LLM InferencevLLM generates response with streaming output
- 05Payment HandoffOrder summary generation and POS integration
Live demo
Try it yourself
Tech stack
Built with
Working on something like this?