← back to projects

AI Restaurant Ordering

A waiter that answers menu and allergen questions from retrieval rather than model memory, on a locally hosted LLM.

Results

What it did

< 500msfirst token streamed
RAGallergens read from the menu, not the model
$0per-token cost — vLLM runs local

The problem

Why it needed building

Restaurant ordering systems are rigid — fixed menus, no dietary intelligence, no upselling. Customers want to ask natural questions like "What's good for someone with gluten allergies?" and get intelligent, context-aware responses. The system needed to run with low latency, handle concurrent conversations, and work with frequently changing menus.

The approach

How it works

Designed a full-stack system with vLLM for local LLM inference (eliminating API costs at scale) combined with a RAG pipeline backed by a vector database for menu retrieval. The conversation engine maintains multi-turn context, understands dietary restrictions, suggests complementary items, and handles payment handoff. WebSocket-based real-time streaming delivers responses as they generate.

  1. User Message
  2. RAG Retrieval
  3. Context Assembly
  4. LLM Inference
  5. Payment Handoff
  1. 01
    User MessageNatural language input via WebSocket connection
  2. 02
    RAG RetrievalVector DB queries menu items, allergens, and pairings
  3. 03
    Context AssemblyConversation history + menu data + dietary preferences
  4. 04
    LLM InferencevLLM generates response with streaming output
  5. 05
    Payment HandoffOrder summary generation and POS integration

Live demo

Try it yourself

Bistro AI● Online
Welcome to Bistro AI! 🍽️ I'm your AI waiter. How can I help you today?

Tech stack

Built with

vLLMRAGFastAPIWebSocketVector DBChromaDBPythonRedis

Working on something like this?