ShadeStudio: System Design
Solution architecture document

Hair color commerce and salon booking platform with natural-language booking

An omnichannel platform where one customer profile powers two journeys: color at home (shade matching, kits, subscriptions) and color with a pro (salon services booked by form, chat or voice).

Prepared by
Kamal Mohan
Role
Tech Lead
Version
1.0, draft for review
Date
October 2026
2 journeysAt home and in salon, one profile
3 ways to bookGuided form, chat, voice
0 double bookingsEnforced in the database
99.9%Target availability, booking path
01

Experience vision

The design starts from the customer, not the services. Every architectural choice below traces back to one of these principles.

One profile, two journeys

Shade, formula, gray coverage, allergies and visit history are shared, so the salon knows what the customer used at home and the app knows what the stylist applied.

Book the way you talk

"Root touch-up Saturday after 4 near me" should work in the app, on WhatsApp or by voice, and always land on the same booking engine.

Confidence before commitment

Quiz, shade comparison and virtual try-on reduce the fear of the wrong color. The agent always reads back slot and price before confirming.

Never lose the thread

A conversation started in chat can finish in the form with fields prefilled, or hand off to a human with full context.

Replenish on rhythm

Subscriptions follow regrowth cycles (typically 4 to 8 weeks) and can be skipped, moved or swapped for a salon visit.

Accessible by default

WCAG 2.2 AA, voice input, large tap targets, regional language support in the conversational layer.

Personas and primary journeys

PersonaGoalJourneyKey touchpoints
At-home colorerCover grays reliably every 6 weeksQuiz → shade match → try-on → buy kit → subscribeWeb, app, email, push
Salon guestQuick professional root serviceAsk in chat → pick slot → pay or use credit → visit → rebookWhatsApp, app, SMS, salon tablet
Hybrid memberMix home kits with periodic salon visitsMembership → perks at both → unified historyAll channels
StylistKnow the guest before they sit downDaily schedule → guest formula card → record service → upsellSalon tablet, POS
Salon managerFill chairs, reduce no-showsRoster → capacity view → waitlist → reportsAdmin console
02

High-level architecture

A layered, API-first design. Channels are thin; the experience layer adapts to each channel; domain services own the business rules; everything talks through events where it does not need an immediate answer.

Channels
Web and PWANext.js, SSR for SEO
Mobile appsReact Native
Chat and voiceIn-app, WhatsApp, SMS, IVR
Salon tabletStylist and kiosk app
Admin consoleRoster, catalog, reports
↓
Edge
CDN and WAFStatic assets, bot and DDoS protection
API gatewayAuth, rate limits, routing
↓
Experience layer
BFFGraphQL for web/app, REST for partners
NL booking agentLLM with tool calling
IdentityOIDC, OTP login, consent
Headless CMSContent, promos, tutorials
Feature flagsExperiments, gradual rollout
↓
Domain services
Catalog and searchProducts, shades, services
Color advisorQuiz, matching, try-on
Cart and checkoutPricing, tax, orders
SubscriptionsAuto-replenish, dunning
Booking engineAvailability, holds, waitlist
Membership and loyaltyTiers, credits, referrals
Customer profileHair history, preferences
NotificationsPush, SMS, email, WhatsApp
Salon operationsLocations, roster, services
InventoryWarehouse and salon stock
↓
Platform and data
Event busKafka, schema registry
PostgreSQLSystem of record
RedisCache, slot holds, sessions
OpenSearchCatalog and location search
Vector storeFAQ and policy retrieval
Object storageImages, try-on assets
WarehouseAnalytics, CDP, ML features
↓
External integrations
PaymentsCards, UPI, wallets
LLM providerBehind a model gateway
MessagingWhatsApp Business, SMS, email
SpeechSpeech-to-text, text-to-speech
ShippingCarriers, tracking
MapsGeocoding, distance
Salon POSIn-store payments

Highlighted boxes form the natural-language booking path.

03

Natural-language booking

The agent is a new front door to the existing booking engine, not a second booking system. The LLM interprets and proposes; deterministic services validate and commit.

What the agent does

  • Understands intent: book, reschedule, cancel, ask, recommend
  • Extracts slots: service, location, date, time window, stylist, party size
  • Fills gaps from profile: home salon, usual service, last stylist
  • Asks one clarifying question at a time when something is missing
  • Calls typed tools, never the database
  • Reads back slot and price, waits for explicit confirmation

Guardrails

  • Tool allow-list with JSON schema validation on every call
  • All writes require a confirmation_token issued after user says yes
  • Prices and availability come only from tool results, never from the model
  • Customer identity taken from the session, never from message text
  • PII redacted before prompts are logged
  • Fallback to form or human after two failed turns or low confidence

Agent tool contract

ToolPurposeType
find_locations(near, radius_km)Salons near a place or the user's saved addressRead
list_services(location_id)Services, duration, price at a salonRead
search_availability(location_id, service_id, from, to, stylist_id?)Open slots in a time windowRead
hold_slot(slot_id)Temporary hold, 5 minute expirySoft write
confirm_booking(hold_id, confirmation_token, payment_ref?)Commit the appointmentWrite
reschedule_booking(booking_id, new_slot_id, confirmation_token)Move an appointmentWrite
cancel_booking(booking_id, confirmation_token)Cancel, applying policyWrite
get_profile()Preferences, last service, membership creditsRead
search_help(query)Retrieve policy and FAQ passagesRead
handoff(reason)Transfer to form or human with contextControl

Agent decision flow

flowchart TD
  A["User message: text or transcribed voice"] --> B["Load session, profile and conversation state"]
  B --> C["LLM: classify intent and extract slots"]
  C --> D{"Booking related?"}
  D -- "No" --> E["Answer from retrieval or route to support"]
  D -- "Yes" --> F{"All required slots filled?"}
  F -- "No" --> G["Ask one clarifying question"]
  G --> A
  F -- "Yes" --> H["Tool: search_availability"]
  H --> I{"Slots found?"}
  I -- "No" --> J["Offer nearest alternatives or waitlist"]
  J --> A
  I -- "Yes" --> K["Tool: hold_slot, then read back slot and price"]
  K --> L{"User confirms?"}
  L -- "No or changes" --> A
  L -- "Yes" --> M["Tool: confirm_booking with confirmation token"]
  M --> N["Send confirmation on same channel"]
  C -. "Low confidence or 2 failed turns" .-> O["Handoff to prefilled form or human"]
  

Figure 1. The model handles language; every state change passes through a validated tool.

Example exchange

User   : Need a root touch-up this Saturday evening, same place as last time
Agent  : [get_profile] → home salon Indiranagar, last service Roots
         [search_availability] → 16:30, 17:15, 18:00
Agent  : Indiranagar has 4:30, 5:15 or 6:00 pm on Saturday. Which works?
User   : 5:15
Agent  : [hold_slot] → hold_81f, expires in 5 min
Agent  : Roots at Indiranagar, Sat 5:15 pm, 45 min, ₹1,800.
         You have one member credit. Use it and confirm?
User   : Yes
Agent  : [confirm_booking] → BK-20931
Agent  : Booked. I'll remind you Friday evening.

3.1 Chat and voice: one agent, two modalities

Chat and voice are adapters around the same agent. A channel adapter turns every input into a text turn tagged with its modality; a response shaper turns the agent's answer into buttons and cards for chat, or a short spoken sentence for voice. Tools, guardrails and the booking engine are identical, so a slot booked by voice behaves exactly like one booked by chat or the form.

flowchart LR
  subgraph CHAT["Chat"]
    C1["In-app chat"]
    C2["WhatsApp / SMS"]
  end
  subgraph VOICE["Voice"]
    V1["In-app mic (WebRTC)"]
    V2["Phone call (telephony / IVR)"]
    V3["Voice activity detection"]
    V4["Streaming speech-to-text"]
    V1 --> V3
    V2 --> V3
    V3 --> V4
  end
  C1 --> AD["Channel adapter: text turn + modality"]
  C2 --> AD
  V4 --> AD
  AD --> AG["NL booking agent"]
  AG --> TL["Booking tools"]
  TL --> BK["Booking engine"]
  AG --> RS["Response shaper"]
  RS -- "chat: text, quick replies, cards" --> CO["Message to chat"]
  RS -- "voice: short sentence" --> TTS["Streaming text-to-speech"]
  TTS --> AO["Audio to caller"]
  

Figure 1a. Two input pipelines converge on one agent; only the edges differ.

AspectChatVoice
Entry pointsIn-app chat, WhatsApp, SMSIn-app mic, phone call to salon number
Reply styleTwo or three lines, can include detailOne short sentence, one question at a time
Offering slotsUp to five as tappable quick repliesAt most two spoken options, then "or another time?"
ConfirmationSummary card, tap or type yesFull read-back of service, salon, day, time and price, then a spoken yes
InterruptionsNot applicableBarge-in: playback stops the moment the caller speaks
Input errorsTypos, handled by the modelMisheard names, dates and numbers; critical entities are repeated back, keypad input as backup
Latency targetFirst token under 1.5 s, streamedAbout 1 s from end of speech to start of reply; a brief acknowledgement covers tool calls
IdentityLogged-in session or verified WhatsApp numberCaller ID matched to profile; one-time code before changing or cancelling
LanguageFree text, mixed Hindi and EnglishLanguage detected in the first turn; speech models selected per language
FallbackPrefilled guided form, or human agentSMS link to prefilled form, or transfer to the salon front desk
After bookingConfirmation in the same threadSpoken confirmation plus SMS or WhatsApp summary (async)

Voice turn sequence

sequenceDiagram
  autonumber
  actor U as Caller
  participant VG as Voice gateway
  participant STT as Speech-to-text
  participant AG as NL agent
  participant BK as Booking service
  participant TTS as Text-to-speech
  participant N as Notifications
  U->>VG: speaks "Roots on Saturday after four"
  VG->>STT: audio stream
  STT--)AG: partial transcripts (streamed)
  STT->>AG: final transcript at end of speech
  AG--)TTS: "One moment, checking Saturday"
  TTS--)U: acknowledgement audio
  AG->>BK: search_availability
  BK-->>AG: 4:30 and 5:15
  AG->>TTS: "I have 4:30 or 5:15. Which do you prefer?"
  TTS--)U: audio (streamed)
  U->>VG: interrupts "5:15"
  VG--)TTS: barge-in, stop playback
  VG->>STT: audio stream
  STT->>AG: "5:15"
  AG->>BK: hold_slot
  AG->>TTS: read back service, salon, time, price
  TTS--)U: audio
  U->>VG: "Yes"
  STT->>AG: confirmation
  AG->>BK: confirm_booking + token
  BK-->>AG: booking id
  AG->>TTS: "You're booked for Saturday 5:15"
  BK--)N: BookingConfirmed (async)
  N--)U: SMS / WhatsApp summary
  

Figure 1b. Dashed arrows are asynchronous or streamed. Speech is streamed in both directions so the caller never waits in silence.

Voice-specific design points

Turn-taking and barge-in

Voice activity detection decides when the caller has finished. If they speak during playback, audio stops at once and the partial reply is discarded from conversation state.

Read-back before commit

The caller cannot see the slot, so the agent always repeats service, salon, day, time and price, and commits only on a clear spoken yes. Anything ambiguous is treated as no.

Filling silence

Tool calls take a few hundred milliseconds. A short acknowledgement is spoken immediately while availability is fetched, keeping the call natural.

Recognition confidence

Low-confidence transcripts trigger a targeted re-ask ("Did you say Saturday the tenth?") rather than a guess. Salon and stylist names are supplied to the recogniser as hint phrases.

Privacy and consent

Callers are told the call is handled by an assistant and may be recorded. Audio is retained briefly for quality review; transcripts are redacted before storage.

Shared conversation state

State is keyed by customer, not channel. A call that drops can resume on WhatsApp with the held slot and collected details intact.

04

Flow diagrams

The four core flows of the platform.

4.1 Booking sequence with slot hold

sequenceDiagram
  autonumber
  actor C as Customer
  participant CH as Channel
  participant AG as NL agent
  participant BK as Booking service
  participant R as Redis
  participant DB as Postgres
  participant PAY as Payments
  participant BUS as Event bus
  C->>CH: "Root touch-up Saturday after 4"
  CH->>AG: message + session
  AG->>BK: search_availability
  BK->>DB: read roster and bookings
  BK-->>AG: open slots
  AG-->>C: offer 3 slots
  C->>AG: picks 5:15
  AG->>BK: hold_slot
  BK->>R: SET hold NX EX 300
  R-->>BK: OK
  BK-->>AG: hold id
  AG-->>C: read back slot and price
  C->>AG: Yes
  AG->>BK: confirm_booking + token
  BK->>PAY: authorize deposit or redeem credit
  PAY-->>BK: authorized
  BK->>DB: insert booking + outbox row in one transaction
  DB-->>BK: committed
  BK->>R: delete hold
  BK-->>AG: booking id
  AG-->>C: confirmation
  DB--)BUS: BookingConfirmed via outbox relay
  

Figure 2. Solid arrows are synchronous. The final dashed arrow is asynchronous: downstream work happens after the customer already has their answer.

4.2 Appointment lifecycle

stateDiagram-v2
  [*] --> Available
  Available --> Held: hold_slot
  Held --> Available: hold expires
  Held --> Confirmed: confirm_booking
  Confirmed --> Confirmed: reschedule
  Confirmed --> Cancelled: cancel
  Confirmed --> CheckedIn: guest arrives
  Confirmed --> NoShow: grace period passes
  CheckedIn --> Completed: service recorded
  Cancelled --> Available: slot released
  NoShow --> [*]
  Completed --> [*]
  

Figure 3. A slot can only be confirmed from a live hold, which is what makes concurrent requests safe.

4.3 Shade match to purchase

flowchart LR
  A["Quiz: current color, gray %, goal, hair history"] --> B["Rules engine: eligible shade families"]
  B --> C["Ranker: top shades by profile similarity"]
  C --> D["Virtual try-on: on-device hair segmentation"]
  D --> E{"Confident?"}
  E -- "No" --> F["Chat with a colorist"]
  F --> C
  E -- "Yes" --> G["Add kit to cart"]
  G --> H{"Subscribe?"}
  H -- "Yes" --> I["Create subscription with cadence"]
  H -- "No" --> J["One-time checkout"]
  I --> K["Order placed"]
  J --> K
  

Figure 4. Rules first for safety (for example, not lifting dark hair to blonde with a deposit-only color), ranking second for personalisation.

4.4 Checkout saga

sequenceDiagram
  autonumber
  participant CO as Checkout
  participant INV as Inventory
  participant PAY as Payments
  participant ORD as Orders
  participant BUS as Event bus
  participant SHIP as Fulfilment
  CO->>INV: reserve stock
  INV-->>CO: reserved
  CO->>PAY: authorize (idempotency key)
  alt authorized
    PAY-->>CO: ok
    CO->>ORD: create order + outbox
    ORD--)BUS: OrderPlaced
    BUS--)SHIP: create shipment
    SHIP--)BUS: ShipmentCreated
    BUS--)PAY: capture payment
  else declined
    PAY-->>CO: declined
    CO->>INV: release stock (compensation)
  end
  

Figure 5. Reserve and authorize are synchronous because the customer is waiting; capture and fulfilment are asynchronous with compensating actions on failure.

05

Asynchronous calls and events

Rule of thumb: if the customer is waiting on the answer, call synchronously with a tight timeout. If they are not, publish an event. This keeps the booking and checkout paths fast and lets side effects fail and retry without affecting the customer.

InteractionModeWhy
Search availability, hold, confirmSyncCustomer needs an immediate, consistent answer
Payment authorizationSyncMust know the result before committing
LLM responseSync, streamedTokens streamed over SSE/WebSocket so replies feel instant
Payment capture, refundsAsyncTriggered by fulfilment or cancellation events; confirmed by webhook
Confirmations and remindersAsyncProvider latency and failures must not block booking
Loyalty credits, referral rewardsAsyncEventually consistent is acceptable
Search index and cache updatesAsyncChange events keep read models fresh
Profile and hair history updateAsyncDerived from completed services and orders
Waitlist offersAsyncReacts to cancellations
Analytics, CDP, ML featuresAsyncStream to the warehouse
Inbound WhatsApp/SMS messagesAsyncWebhook is acknowledged at once, message queued for the agent

Event fan-out

flowchart LR
  BK["Booking service"] --> OB[("Outbox table")]
  CO["Checkout / Orders"] --> OB
  SUB["Subscriptions"] --> OB
  OB --> RL["Outbox relay (CDC)"]
  RL --> K{{"Event bus"}}
  K --> N["Notifications"]
  K --> L["Loyalty"]
  K --> P["Profile"]
  K --> S["Search indexer"]
  K --> W["Waitlist"]
  K --> F["Fulfilment"]
  K --> A["Analytics and CDP"]
  N -. "after max retries" .-> DLQ[("Dead-letter queue")]
  F -. "after max retries" .-> DLQ
  

Figure 6. The transactional outbox guarantees an event is published if and only if the business transaction committed.

Event catalog

EventProducerConsumers and effect
BookingConfirmedBookingNotifications (confirm, schedule reminders), Loyalty (hold credit), Analytics
BookingRescheduledBookingNotifications (update reminders), Waitlist (old slot freed)
BookingCancelledBookingPayments (refund per policy), Waitlist (offer slot), Loyalty (return credit)
ServiceCompletedSalon opsProfile (formula and history), Loyalty (earn points), Notifications (review and rebook prompt)
NoShowRecordedBookingPayments (fee if applicable), Profile (reliability signal)
OrderPlacedOrdersFulfilment, Inventory (decrement), Notifications, Loyalty
ShipmentCreated / DeliveredFulfilmentPayments (capture), Notifications (tracking), Subscriptions (reset cadence clock)
SubscriptionRenewalDueSubscriptionsNotifications (3-day heads-up), Checkout (create renewal order)
PaymentFailedPaymentsSubscriptions (dunning schedule), Notifications
CatalogChangedCatalogSearch indexer, CDN cache purge
ConversationEndedNL agentAnalytics (containment, drop-off), Evaluation pipeline

Reminder and subscription timing

sequenceDiagram
  participant BUS as Event bus
  participant SCH as Scheduler
  participant N as Notifications
  participant PR as Messaging provider
  actor C as Customer
  BUS--)SCH: BookingConfirmed
  SCH->>SCH: schedule T-24h and T-2h jobs
  Note over SCH: time passes
  SCH--)N: ReminderDue (T-24h)
  N->>PR: send WhatsApp template
  PR--)N: delivery webhook
  C--)PR: replies "reschedule"
  PR--)N: inbound webhook
  N--)BUS: InboundMessage to NL agent queue
  

Figure 7. Reminders are delayed jobs; a reply to a reminder re-enters the NL agent, so rescheduling works straight from the notification.

Reliability patterns

Transactional outbox

Business row and event row are written in one database transaction; a relay (change data capture) publishes to the bus. No dual-write gap.

Idempotent consumers

Delivery is at-least-once. Each consumer stores processed event IDs, so a redelivered BookingConfirmed does not send two messages or award double points.

Retries and dead letters

Exponential backoff with jitter, capped attempts, then a dead-letter queue with alerts and a replay tool.

Ordering

Events are partitioned by aggregate ID (booking, order), so updates for one booking are processed in order while different bookings run in parallel.

Sagas with compensation

Multi-step flows (checkout, cancellation refunds) are orchestrated; each step has an undo action instead of a distributed transaction.

Webhooks

Payment and messaging webhooks are signature-verified, acknowledged immediately, queued, then processed idempotently.

Scheduled and background jobs

JobCadencePurpose
Subscription renewal scanHourlyEmit SubscriptionRenewalDue for upcoming cycles
Dunning retriesDay 1, 3, 5, 7Retry failed renewals, then pause
No-show sweeperEvery 5 minMark appointments past the grace period
Availability materialiserNightly and on roster changePrecompute slots for the next 60 days
Payment reconciliationDailyMatch provider settlements against orders and bookings
Abandoned cart and quiz nudgesHourlyConsent-aware re-engagement
Conversation evaluationDailyScore sampled transcripts for accuracy and safety
06

Domain services

Each service owns its data and exposes an API plus events. They start as modules in one deployable and are split out when scale or team boundaries require it.

ServiceResponsibilitiesOwnsStore
Catalog and searchProducts, shades, salon services, pricing, faceted searchProduct, Shade, ServicePostgres, OpenSearch
Color advisorQuiz, rules, shade ranking, try-on assets, colorist chatQuizResult, RecommendationPostgres, object storage
Cart and checkoutCart, promotions, tax, payment orchestration, ordersCart, Order, PaymentRedis (cart), Postgres
SubscriptionsPlans, cadence, skip/swap, renewals, dunningSubscriptionPostgres
Booking engineAvailability, holds, appointments, waitlist, policiesSlot, Booking, WaitlistPostgres, Redis
Salon operationsLocations, hours, stylists, roster, skillsLocation, Stylist, ShiftPostgres
Membership and loyaltyTiers, credits, points ledger, referralsMembership, LedgerEntryPostgres
Customer profileIdentity link, hair profile, formulas, consentCustomer, HairProfilePostgres
NotificationsTemplates, channel preference, scheduling, delivery trackingMessage, TemplatePostgres, queue
NL booking agentConversation state, prompt assembly, tool execution, handoffConversation, TurnRedis, Postgres
07

Data model

Core entities for the booking and commerce paths.

erDiagram
  CUSTOMER ||--o| HAIR_PROFILE : has
  CUSTOMER ||--o{ BOOKING : makes
  CUSTOMER ||--o{ ORDER : places
  CUSTOMER ||--o{ SUBSCRIPTION : holds
  CUSTOMER ||--o| MEMBERSHIP : has
  CUSTOMER ||--o{ CONVERSATION : starts
  LOCATION ||--o{ STYLIST : employs
  LOCATION ||--o{ BOOKING : hosts
  STYLIST ||--o{ SHIFT : works
  STYLIST ||--o{ BOOKING : serves
  SERVICE ||--o{ BOOKING : booked_as
  ORDER ||--|{ ORDER_ITEM : contains
  PRODUCT ||--o{ ORDER_ITEM : sold_as
  PRODUCT ||--o{ SUBSCRIPTION : replenishes
  CUSTOMER {
    uuid id PK
    string phone
    string email
    string locale
  }
  HAIR_PROFILE {
    uuid customer_id FK
    string natural_level
    int gray_percent
    string current_shade
    string allergies
  }
  BOOKING {
    uuid id PK
    uuid customer_id FK
    uuid stylist_id FK
    uuid service_id FK
    uuid location_id FK
    datetime starts_at
    datetime ends_at
    string status
    string channel
  }
  SHIFT {
    uuid id PK
    uuid stylist_id FK
    datetime starts_at
    datetime ends_at
  }
  SERVICE {
    uuid id PK
    string name
    int duration_min
    int price_minor
  }
  ORDER {
    uuid id PK
    uuid customer_id FK
    int total_minor
    string status
  }
  SUBSCRIPTION {
    uuid id PK
    uuid customer_id FK
    uuid product_id FK
    int cadence_weeks
    date next_run
    string status
  }
  CONVERSATION {
    uuid id PK
    uuid customer_id FK
    string channel
    string outcome
  }
  

Figure 8. Money is stored in minor units; timestamps in UTC with the salon's time zone stored on the location.

7.1 Database tables

PostgreSQL, one schema per module so a module can be lifted into its own database later. All tables carry id uuid primary keys, created_at and updated_at; these are omitted below for brevity.

TableKey columnsKeys, constraints, indexesOwner
customerphone, email, full_name, locale, statusUnique (phone), unique (email)Profile
hair_profilecustomer_id, natural_level, gray_percent, current_shade_id, texture, allergies, last_colored_atUnique (customer_id); FK customerProfile
consent_recordcustomer_id, purpose, channel, granted, source, recorded_atAppend-only; index (customer_id, purpose)Profile
locationname, address, geo (point), time_zone, opening_hours (jsonb), statusGiST index (geo)Salon ops
stylistlocation_id, display_name, level, activeFK location; index (location_id, active)Salon ops
servicename, category, duration_min, buffer_min, price_minor, currency, activeUnique (name)Catalog
stylist_servicestylist_id, service_idComposite PK; which stylist can perform which serviceSalon ops
shiftstylist_id, starts_at, ends_at, kind (work, break, leave)Exclusion on overlapping shifts per stylistSalon ops
slotlocation_id, stylist_id, service_id, starts_at, ends_at, stateMaterialised read model; index (location_id, service_id, starts_at) where state = 'open'Booking
bookingcustomer_id, location_id, stylist_id, service_id, starts_at, ends_at, status, channel, price_minor, payment_id, idempotency_key, versionExclusion constraint on (stylist_id, time range); unique (idempotency_key); index (customer_id, starts_at)Booking
booking_eventbooking_id, from_status, to_status, actor_type, actor_id, reasonAppend-only status history; index (booking_id)Booking
waitlist_entrycustomer_id, location_id, service_id, window_start, window_end, statusIndex (location_id, service_id, window_start)Booking
productsku, name, type, shade_id, price_minor, activeUnique (sku)Catalog
shadecode, name, family, level, tone, gray_coverageUnique (code); index (family, level)Catalog
inventorysku, site_id, on_hand, reservedComposite PK (sku, site_id); check (on_hand >= reserved)Inventory
orderscustomer_id, status, subtotal_minor, tax_minor, total_minor, shipping_address (jsonb), subscription_id, idempotency_keyUnique (idempotency_key); index (customer_id, created_at)Checkout
order_itemorder_id, sku, qty, unit_price_minorFK orders; index (order_id)Checkout
paymentcustomer_id, purpose (order, booking), ref_id, provider, provider_ref, amount_minor, statusUnique (provider, provider_ref); no card data storedCheckout
subscriptioncustomer_id, sku, cadence_weeks, next_run, status, payment_token_ref, retry_countPartial index (next_run) where status = 'active'Subscriptions
membershipcustomer_id, tier, started_at, renews_at, service_credits, statusUnique (customer_id)Loyalty
loyalty_ledgercustomer_id, delta, reason, source_event_idAppend-only; unique (source_event_id) makes awards idempotentLoyalty
conversationcustomer_id, channel, modality (chat, voice), language, started_at, ended_at, outcomeIndex (customer_id, started_at)NL agent
conversation_turnconversation_id, seq, role, text_redacted, tool_name, tool_args (jsonb), latency_ms, stt_confidenceUnique (conversation_id, seq); partitioned by monthNL agent
notificationcustomer_id, template, channel, send_at, status, provider_ref, source_event_idUnique (source_event_id, template); index (send_at) where status = 'scheduled'Notifications
outbox_eventaggregate_type, aggregate_id, event_type, payload (jsonb), published_atPartial index where published_at is nullEvery module
processed_eventconsumer, event_id, processed_atComposite PK (consumer, event_id)Every consumer
staff_useremail, role, location_id, mfa_enrolled, activeUnique (email)Identity
audit_logactor_type, actor_id, action, entity, entity_id, ip, atAppend-only; partitioned by monthPlatform

Core DDL

CREATE TABLE booking (
  id               uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  customer_id      uuid NOT NULL REFERENCES customer(id),
  location_id      uuid NOT NULL REFERENCES location(id),
  stylist_id       uuid NOT NULL REFERENCES stylist(id),
  service_id       uuid NOT NULL REFERENCES service(id),
  starts_at        timestamptz NOT NULL,
  ends_at          timestamptz NOT NULL,
  status           text NOT NULL CHECK (status IN
                     ('confirmed','checked_in','completed','cancelled','no_show')),
  channel          text NOT NULL CHECK (channel IN
                     ('web','app','chat','whatsapp','voice','salon')),
  price_minor      integer NOT NULL CHECK (price_minor >= 0),
  payment_id       uuid REFERENCES payment(id),
  idempotency_key  text NOT NULL UNIQUE,
  version          integer NOT NULL DEFAULT 1,        -- optimistic locking
  created_at       timestamptz NOT NULL DEFAULT now(),
  updated_at       timestamptz NOT NULL DEFAULT now(),
  CHECK (ends_at > starts_at)
);
CREATE INDEX booking_customer_idx ON booking (customer_id, starts_at DESC);
CREATE INDEX booking_location_day_idx ON booking (location_id, starts_at);

CREATE TABLE outbox_event (
  id              uuid PRIMARY KEY DEFAULT gen_random_uuid(),
  aggregate_type  text NOT NULL,
  aggregate_id    uuid NOT NULL,                      -- Kafka partition key
  event_type      text NOT NULL,
  payload         jsonb NOT NULL,
  created_at      timestamptz NOT NULL DEFAULT now(),
  published_at    timestamptz
);
CREATE INDEX outbox_unpublished_idx ON outbox_event (created_at)
  WHERE published_at IS NULL;

CREATE TABLE processed_event (
  consumer      text NOT NULL,
  event_id      uuid NOT NULL,
  processed_at  timestamptz NOT NULL DEFAULT now(),
  PRIMARY KEY (consumer, event_id)
);

Data management

Partitioning and retention

High-volume append-only tables (conversation_turn, audit_log, notification) are partitioned by month. Old partitions are archived to object storage and dropped; published outbox rows are purged after 7 days.

Read scaling

Availability search reads the slot read model through Redis and read replicas. Writes and the confirm path always use the primary.

Sensitive columns

Phone, email, address and allergies are encrypted at column level with keys in a key management service. Conversation text is stored only after redaction.

Migrations

Versioned, backward-compatible migrations (expand, migrate, contract) so deploys never need downtime or a locked table.

Preventing double booking

Two layers. Redis gives a fast, expiring hold so two people do not walk through checkout for the same chair. The database gives the hard guarantee with an exclusion constraint, so even if Redis fails, overlapping appointments cannot be committed.

-- Fast path: hold for 5 minutes, only if nobody else holds it
SET hold:{stylist_id}:{starts_at} {customer_id} NX EX 300

-- Hard guarantee: no overlapping active bookings per stylist
ALTER TABLE booking ADD CONSTRAINT no_overlap
  EXCLUDE USING gist (
    stylist_id WITH =,
    tstzrange(starts_at, ends_at) WITH &&
  ) WHERE (status IN ('confirmed','checked_in'));
08

API contracts

The guided form and the NL agent call the same endpoints. All writes accept an Idempotency-Key header.

Method and pathPurpose
GET /v1/locations?near=lat,lng&radius_km=Find salons
GET /v1/locations/{id}/servicesServices and prices
GET /v1/availability?location_id=&service_id=&from=&to=Open slots
POST /v1/holdsHold a slot
POST /v1/bookingsConfirm from a hold
PATCH /v1/bookings/{id}Reschedule
DELETE /v1/bookings/{id}Cancel
POST /v1/conversations/{id}/messagesSend a message to the agent (SSE response stream)
POST /v1/quiz/resultsSubmit quiz, get shade recommendations
POST /v1/checkoutPlace order
POST /v1/subscriptions · PATCH /v1/subscriptions/{id}Create, skip, swap, pause
POST /v1/webhooks/{provider}Payment and messaging callbacks
POST /v1/bookings
Idempotency-Key: 7c1e...

{ "hold_id": "hold_81f", "payment": { "type": "member_credit" }, "channel": "whatsapp" }

201 Created
{ "id": "BK-20931", "status": "confirmed",
  "starts_at": "2026-10-10T17:15:00+05:30",
  "location": "Indiranagar", "service": "Roots", "price_minor": 180000 }

409 Conflict  → { "code": "HOLD_EXPIRED", "alternatives": [ ... ] }
09

Tech stack

Chosen for one language across the stack (TypeScript), managed services where they remove operational load, and open standards at the boundaries. Named products are reasonable defaults, each replaceable behind an interface.

LayerChoiceWhy
WebNext.js (React, TypeScript)Server rendering for SEO-heavy catalog pages, PWA support
MobileReact NativeShared code and skills with web; native modules for camera
Design systemShared component library, design tokensConsistent experience across web, app and tablet
API layerGraphQL BFF, REST for partners and webhooksChannel-shaped responses without chatty clients
BackendNode.js with NestJS, modular monolithClear module boundaries, easy to split later
ConversationalLLM with tool calling behind a model gatewayProvider-agnostic, central logging, cost and safety controls
Retrievalpgvector in PostgresFAQ and policy grounding without another datastore
VoiceStreaming speech-to-text and text-to-speech APIsVoice is a thin adapter over the same agent
Virtual try-onOn-device hair segmentation (MediaPipe or similar)Low latency, photos never leave the phone
Primary databasePostgreSQLTransactions, exclusion constraints for bookings
Cache and holdsRedisExpiring keys, sessions, rate limits
SearchOpenSearchFaceted product search, geo search for salons
EventsKafka (managed) with schema registry; SQS for simple job queuesDurable, ordered, replayable
SchedulingTemporal or a delayed-job queueReminders, dunning and sagas with retries built in
PaymentsRazorpay or Stripe (cards, UPI, wallets)Tokenised payments, no card data on our servers
MessagingWhatsApp Business API, SMS and email providersMeet customers where they already are
IdentityOIDC provider with OTP and social loginPasswordless on mobile, single identity across channels
CMSHeadless CMSMarketing ships content without engineering releases
CloudAWS: containers on EKS or ECS, RDS, ElastiCache, S3, CloudFrontManaged building blocks, multi-zone
DeliveryGitHub Actions, Terraform, blue-green deploys, feature flagsSmall, safe, reversible releases
ObservabilityOpenTelemetry, Prometheus, Grafana, centralised logsOne trace from chat message to database commit
AnalyticsEvent stream to a warehouse, dbt, product analyticsFunnel, cohort and conversation metrics
10

Non-functional requirements

Figures are design targets and planning assumptions for a mid-size launch, to be validated with load tests.

Performance

  • Availability search p95 under 300 ms (precomputed slots plus cache)
  • Booking confirm p95 under 800 ms
  • First agent token under 1.5 s, streamed
  • Catalog pages LCP under 2.5 s on 4G

Scale assumptions

  • 500k monthly active users, 100 salons
  • About 10k bookings and 15k orders per day
  • Read-heavy: roughly 100 availability reads per booking
  • Peaks: weekends, festive season, campaign launches

Availability and resilience

  • 99.9% for booking and checkout; multi-zone deployment
  • If the LLM is down, chat degrades to the guided form
  • If Redis is down, the database constraint still protects bookings
  • Circuit breakers and timeouts on every external call
  • RPO 5 min, RTO 1 hour

Security and privacy

  • OAuth 2.0 / OIDC, short-lived tokens, role-based access for staff
  • Payment data tokenised by the provider (reduced PCI scope)
  • Encryption in transit and at rest; secrets in a vault
  • Consent ledger; compliance with India's DPDP Act and GDPR where applicable
  • Prompt-injection defences: tool allow-list, server-side identity, output validation

Observability

  • Distributed trace per request, including tool calls
  • Business metrics: booking conversion, hold expiry rate, no-show rate
  • Agent metrics: containment, turns to book, fallback rate, tool error rate
  • Consumer lag and dead-letter alerts

Quality

  • Contract tests between modules and for events
  • Golden-conversation regression suite run on every prompt or model change
  • Load tests on availability and confirm paths
  • Accessibility audits in CI
11

Security architecture

Defence in depth across four trust zones. The conversational layer is treated as untrusted input, the same as any public form.

flowchart LR
  subgraph Z1["Zone 1: public"]
    U["Customers and callers"]
    P["Provider webhooks"]
  end
  subgraph Z2["Zone 2: edge"]
    W["CDN, WAF, bot control"]
    G["API gateway: token check, rate limit"]
  end
  subgraph Z3["Zone 3: application (private subnets)"]
    B["BFF and NL agent"]
    S["Domain modules"]
    M["Model gateway: redaction, allow-list"]
  end
  subgraph Z4["Zone 4: data (no internet route)"]
    D[("Postgres, Redis, Kafka")]
    K["Key management and secrets"]
  end
  U --> W
  P --> W
  W --> G
  G --> B
  B --> S
  B --> M
  S --> D
  S --> K
  M --> L["LLM provider (egress allow-list)"]
  

Figure 9. Each boundary re-validates identity and input; nothing in zone 4 is reachable from the internet.

Identity and access

ActorAuthenticationAuthorisation
Customer (web, app)OIDC with phone OTP or social login; short-lived access token, rotating refresh tokenCan only read and change own profile, orders and bookings (ownership check in every query)
Customer (WhatsApp)Verified sender number mapped to profileSame as above; step-up OTP for cancel, reschedule or address change
Caller (voice)Caller ID match; treated as low assuranceMay book; OTP required before changing or cancelling an existing booking
StylistStaff SSO with MFA on the salon tabletOwn schedule and guest formula cards at own location only
Salon managerStaff SSO with MFARoster, bookings and reports for assigned locations
Support agentStaff SSO with MFACustomer lookup with masked contact fields; every view audited
Service to serviceMutual TLS with workload identityPer-service scopes; least-privilege database roles per module
NL agentRuns with the customer's session, never a super-user credentialTool allow-list; customer ID injected server-side, not from model output

Threats and controls

ThreatControl
Account takeover via OTP abuseOTP rate limits per number and device, attempt lockout, SIM-swap risk signals, device binding
Reading another customer's booking (IDOR)Ownership enforced in the data layer; opaque IDs; automated authorisation tests
Prompt injection ("ignore rules, cancel all bookings")Model cannot widen its own permissions: tools are scoped to the session customer, writes need a server-issued confirmation token, arguments are schema-validated
Model leaking data or inventing pricesOnly tool results are shown as facts; output filter for PII and off-policy content; retrieval limited to public help content
Slot hoarding and booking botsHold limits per customer, bot detection at the edge, deposit for repeat no-shows
Forged payment or messaging webhooksSignature verification, timestamp tolerance, replay protection by event ID, source IP allow-list
Payment fraud and card testingProvider-hosted fields and tokenisation, 3-D Secure, velocity rules, idempotency keys
Injection and cross-site attacksParameterised queries, input validation, content security policy, same-site cookies, CSRF tokens
Voice spoofing or recorded "yes"Caller ID is never sufficient for sensitive changes; OTP step-up; confirmation phrased with a changing detail
Insider misuseRole-based access, masked fields, just-in-time production access, immutable audit log
Supply-chain compromiseDependency and container scanning in CI, signed images, pinned versions, software bill of materials
Denial of service and LLM cost abusePer-user and per-IP rate limits, token budgets per conversation, autoscaling with caps

Data protection

ClassExamplesHandling
RestrictedPayment tokens, OTP secrets, credentialsNever logged; held by provider or secrets vault; no card numbers stored
Sensitive personalAllergies, hair and scalp notes, voice recordings, try-on photosColumn encryption; explicit consent; photos processed on device; recordings kept 30 days
PersonalName, phone, email, address, booking historyEncrypted at rest, masked in support tools and logs, redacted before reaching the LLM where not needed
InternalRoster, pricing rules, sales figuresRole-based access
PublicCatalog, salon addresses, help contentCached at the CDN

Encryption and secrets

  • TLS 1.2 or higher everywhere, mutual TLS inside the cluster
  • AES-256 at rest with managed keys, rotated yearly
  • Secrets in a vault, injected at runtime, never in code or images

Privacy and compliance

  • DPDP Act (India): consent ledger, purpose limitation, erase and export on request
  • GDPR alignment where EU residents are served
  • PCI DSS scope reduced to SAQ A by using hosted payment fields
  • LLM provider contract: no training on customer data, regional processing

Secure delivery

  • Static analysis, dependency and secret scanning on every pull request
  • Infrastructure as code with policy checks
  • Annual penetration test, plus red-team prompts against the agent each release

Detection and response

  • Central security log with anomaly alerts (OTP spikes, unusual cancellations, bulk lookups)
  • Immutable audit trail for staff and agent write actions
  • Incident runbooks with breach notification timelines
12

Metrics and observability

Three layers of measurement: is the system healthy, is the agent doing its job, and is the business outcome improving. Targets are proposed starting points to tune after launch.

flowchart LR
  A["Services, agent, workers"] -- "traces, metrics, logs" --> O["OpenTelemetry collector"]
  O --> P["Prometheus: metrics"]
  O --> T["Trace store"]
  O --> G["Log store (PII redacted)"]
  P --> D["Grafana dashboards"]
  T --> D
  G --> D
  P --> AL["Alert manager"]
  AL --> PG["On-call paging"]
  A -- "business events" --> K{{"Event bus"}}
  K --> W[("Warehouse")]
  W --> BI["Product and business dashboards"]
  W --> EV["Agent evaluation pipeline"]
  

Figure 10. Operational telemetry and business events travel separate paths so analytics load never affects alerting.

Service level objectives

JourneyIndicatorObjective (30 days)Page when
Availability searchSuccessful responses under 300 ms99% of requestsError budget burning 10x for 5 min
Booking confirmSuccess rate, excluding genuine conflicts99.9%Below 99% for 5 min
Booking confirmLatency p95Under 800 msAbove 1.5 s for 10 min
CheckoutOrder placement success99.9%Payment failure rate doubles against baseline
Chat agentTime to first token p95Under 1.5 sAbove 3 s for 10 min
Voice agentEnd of speech to first audio p95Under 1.2 sAbove 2 s for 5 min
NotificationsConfirmation sent within 60 s of booking99%Consumer lag above 2 min
Data integrityDouble bookingsZeroAny occurrence

Conversational metrics

MetricDefinitionTarget
Containment rateConversations completed without form or human handoffChat 70%, voice 55%
Booking completionBooking intents that end in a confirmed bookingAbove 60%
Turns to bookMedian customer turns from intent to confirmation4 or fewer
Wrong-booking rateBookings changed or cancelled within 10 min citing an errorBelow 0.5%
Slot-extraction accuracyService, date, time and location correct on the golden setAbove 95%
Tool-call validityTool calls passing schema validation first timeAbove 99%
Fallback rateHandoffs by reason (low confidence, repeated failure, user request)Trend down
Word error rate (voice)Speech recognition errors on sampled calls, per languageBelow 12%
Barge-in rate (voice)Replies interrupted by the caller; a proxy for replies that are too longBelow 20%
Cost per conversationLLM tokens plus speech minutesBudget tracked weekly
Safety violationsOff-policy or leaked content found by evaluationZero tolerated

Business and product metrics

AreaMetricWhy it matters
BookingSearch-to-book conversion by channelShows whether NL booking beats the form
BookingHold expiry rateHigh values mean friction at confirmation or payment
SalonChair utilisation, no-show rate, rebook rate within 8 weeksCore salon economics
CommerceQuiz completion, quiz-to-purchase conversion, shade return rateQuality of shade matching
SubscriptionsRetention at 3 and 6 months, skip rate, recovery after failed paymentRecurring revenue health
MembershipCredit redemption, member versus non-member visit frequencyValue of the membership
Cross-journeyShare of customers using both home kits and salon in 6 monthsThe thesis of the platform
ExperienceCSAT after visit and after conversation, complaint rateCustomer voice

Platform and async metrics

Golden signals per service

  • Request rate, error rate, latency (p50, p95, p99), saturation
  • Database connections, slow queries, replica lag
  • Redis hit ratio and hold-key count

Event pipeline

  • Outbox backlog and age of oldest unpublished event
  • Consumer lag per group, retry counts
  • Dead-letter queue depth (alert on any growth)
  • Scheduled job success and duration

External dependencies

  • LLM, speech, payment and messaging latency and error rate
  • Circuit-breaker state changes
  • Message delivery and read rates by channel

Tracing and logs

  • One trace ID from channel through agent, tool calls and database commit
  • Structured logs with customer ID hashed and PII redacted
  • Conversation ID attached to every span for replay
13

Exception handling

Every failure is classified once, close to where it happens, and then handled by rule: retry it, compensate for it, degrade around it, or tell the customer plainly. No layer swallows an error silently and no raw exception reaches a customer.

flowchart TD
  E["Exception raised"] --> C{"Classify"}
  C -- "Validation or business rule" --> V["Return 4xx with error code, no retry"]
  C -- "Transient: timeout, 5xx, throttled" --> R{"Operation idempotent?"}
  R -- "Yes" --> T["Retry with backoff and jitter, max 3"]
  R -- "No" --> F
  T -- "Recovered" --> OK["Continue"]
  T -- "Exhausted" --> F{"Fallback available?"}
  F -- "Yes" --> D["Degrade: cached data, guided form, queue for later"]
  F -- "No" --> X["Fail the request, compensate completed steps"]
  C -- "Unexpected bug" --> X
  V --> L["Log with trace ID, emit error metric"]
  D --> L
  X --> L
  L --> M["Customer message mapped from error code"]
  

Figure 11. One decision path for every exception, sync or async.

Error taxonomy

ClassExamplesHTTPRetryHandling
ValidationMissing field, bad date, unknown service400, 422NoField-level errors returned; agent re-asks for that slot only
AuthenticationExpired token, failed OTP401After re-authSilent token refresh once, then login prompt
AuthorisationBooking belongs to someone else403, 404NoGeneric not-found to avoid leaking existence; security log entry
Business conflictSlot taken, hold expired, credit already used409NoReturn alternatives in the response body
Rate limitToo many OTPs or holds429After Retry-AfterClient backs off; agent explains the wait
Transient dependencyPayment timeout, LLM 5xx, database failover502, 503, 504Yes, if idempotentBackoff, circuit breaker, fallback
UnexpectedNull reference, unhandled state500NoGlobal handler; generic message; alert on rate

Standard error response

All APIs return the same envelope (RFC 9457 problem details). The code is stable and machine-readable; channels map it to their own wording, so the API never dictates customer-facing text.

HTTP/1.1 409 Conflict
Content-Type: application/problem+json

{
  "type":      "https://api.shadestudio.example/errors/hold-expired",
  "title":     "Slot hold has expired",
  "status":    409,
  "code":      "HOLD_EXPIRED",
  "detail":    "The hold on Saturday 17:15 expired before confirmation.",
  "retryable": false,
  "trace_id":  "4bf92f3577b34da6a3ce929d0e0e4736",
  "alternatives": [ { "slot_id": "sl_7a2", "starts_at": "2026-10-10T18:00:00+05:30" } ]
}

Error code catalogue

CodeMeaningChat responseVoice response
SLOT_UNAVAILABLESlot was taken by someone elseOffer the next three slots as quick replies"That time just went. I have 6 pm instead, shall I take it?"
HOLD_EXPIREDCustomer took longer than 5 minutesRe-check availability and re-hold automatically if still freeSame, silently; only mention it if the slot is gone
PAYMENT_DECLINEDBank refused the paymentKeep the hold; offer another method or pay at salonSend a payment link by SMS and keep the hold
PAYMENT_PENDINGProvider timed out; outcome unknown"Checking with your bank"; resolve by webhook or status pollTell the caller a confirmation message will follow
CREDIT_INSUFFICIENTNo membership credit leftShow price and ask to pay insteadRead the price and ask to continue
OUTSIDE_POLICYCancel or reschedule inside the cut-off windowExplain the fee and ask for confirmationSame, with the fee read aloud
VERIFICATION_REQUIREDStep-up OTP neededSend OTP and ask for itSend OTP by SMS and ask the caller to read it
OUT_OF_STOCKShade kit unavailableOffer notify-me or the closest shadeNot applicable
AGENT_UNAVAILABLELLM or speech provider is downOpen the guided form with fields prefilledSwitch to keypad menu or transfer to the salon
INTERNAL_ERRORUnexpected failureApologise, give a reference ID, offer a humanApologise and transfer to the front desk

Handling by layer

LayerResponsibility
Client (web, app)Error boundaries per screen, offline detection, retry button with the same idempotency key, never shows stack traces
API gatewayRejects malformed and unauthenticated requests early; uniform 429 and 503 responses
BFFPartial responses: a failed recommendations call does not break the page; maps codes to localised messages
NL agentTreats tool errors as data: reads the code, chooses to re-ask, offer alternatives or hand off; invalid model output is repaired once, then falls back
Domain modulesTyped domain exceptions; a single global handler converts them to the standard envelope; transactions roll back as a unit
Integration adaptersTimeouts on every call, retry policy, circuit breaker, translation of provider errors into internal codes
Async workersRetry with backoff, then dead-letter queue; poison messages never block a partition
DatabaseConstraint violations mapped to business codes (exclusion violation becomes SLOT_UNAVAILABLE); optimistic-lock conflicts retried once

Resilience policies

DependencyTimeoutRetriesCircuit breakerFallback
Payment provider8 sNone on authorise; status check insteadOpens at 50% failures over 20 callsPay at salon for bookings; hold order for commerce
LLM provider10 s, 2 s to first token1, then secondary modelOpens at 30% failuresGuided form with prefilled fields
Speech-to-text and text-to-speech3 s1Opens at 30% failuresKeypad menu or transfer to salon
Messaging provider5 s5 with backoff (async)Per channelNext channel in preference order: WhatsApp, SMS, email
Redis100 ms1YesSkip the hold, rely on the database constraint
Search500 ms1YesDatabase query for a reduced result set
Postgres2 s statement timeout1 on serialisation or failover errorsNoFail fast with 503 and Retry-After

Cases that need special care

Payment outcome unknown

A timeout is not a decline. The booking is kept in a pending state, the outcome is resolved by webhook or a status query, and the hold is extended meanwhile. The same idempotency key prevents a double charge on retry.

Paid but booking failed

If payment succeeds and the booking insert fails (for example the slot was lost), the saga issues an automatic void or refund and offers alternatives. A reconciliation job catches anything the saga missed.

Dead-letter handling

Each dead-lettered event raises an alert with its error and trace ID. An operator tool can inspect, fix and replay it; consumers are idempotent so replay is safe.

Model misbehaviour

Malformed tool arguments, a tool not on the allow-list or a reply that fails the output filter are treated as exceptions: one repair attempt, then handoff. These are counted and reviewed in evaluation.

Dropped voice call

Conversation state and any held slot survive the drop. The customer gets an SMS link to finish, or can call back and resume where they stopped.

Partial outage

Feature flags act as kill switches for the agent, try-on, recommendations and promotions, so the core paths (browse, book, pay) stay up while a non-critical part is disabled.

14

Logging and log management

Logs are structured, correlated, redacted at source and kept only as long as they are useful. Application logs explain behaviour; audit logs prove who did what; they are stored and governed separately.

flowchart LR
  A["Services, agent, workers: JSON to stdout"] --> C["Collector agent per node"]
  E["Edge: CDN, WAF, gateway logs"] --> C
  C --> P["Processor: redact PII, enrich, sample"]
  P --> H[("Hot store: searchable, 14 days")]
  P --> W[("Warm store: 90 days")]
  W --> X[("Archive: object storage, 1 year")]
  H --> Q["Search and dashboards"]
  H --> AL["Log-based alerts"]
  AU["Audit events"] --> AS[("Audit store: write-once, 7 years")]
  P --> SI["Security monitoring"]
  AS --> SI
  

Figure 12. Redaction happens in the pipeline as a second line of defence; the first is not logging sensitive values at all.

Log types

TypeContentsRetentionAccess
ApplicationRequest handling, business decisions, errors with stack traces14 days hot, 90 days warmEngineering
AccessMethod, route, status, latency, caller at gateway and CDN90 daysEngineering, security
AuditStaff and agent write actions, data views by support, permission changes7 years, write-onceSecurity, compliance
SecurityLogin attempts, OTP events, WAF blocks, authorisation denials1 yearSecurity
ConversationRedacted turns, tool calls, latency, recognition confidence90 days, then anonymised aggregatesConversation design team, restricted
Voice recordingsCall audio, with consent30 daysQuality reviewers only
IntegrationOutbound calls and webhooks: provider, status, latency (no payloads with card or personal data)90 daysEngineering

Log levels

LevelUsed forExampleProduction
ERRORA request or job failed and needs attentionBooking insert failed after payment authorisedAlways on; feeds alerts
WARNRecovered or degraded, worth watchingRetry succeeded, circuit opened, fallback usedAlways on
INFOBusiness milestones and state changesBooking confirmed, order placed, handoff triggeredAlways on
DEBUGDiagnostic detailTool arguments, query timingsOff; enabled per service or per trace by flag

Expected business outcomes such as SLOT_UNAVAILABLE are logged at INFO, not ERROR, so error rates reflect real faults.

Standard log record

{
  "ts":              "2026-10-10T11:45:02.318Z",
  "level":           "ERROR",
  "service":         "booking",
  "env":             "prod",
  "version":         "1.14.2",
  "trace_id":        "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id":         "00f067aa0ba902b7",
  "conversation_id": "cv_5521",
  "customer_ref":    "c_9f3a...",          // hashed, never phone or email
  "channel":         "voice",
  "event":           "booking.confirm.failed",
  "error_code":      "SLOT_UNAVAILABLE",
  "error_class":     "business_conflict",
  "retryable":       false,
  "duration_ms":     212,
  "message":         "Exclusion constraint rejected overlapping booking"
}

Management practices

Correlation

A trace ID is created at the edge and passed through HTTP headers, event headers and tool calls. One search shows a chat message, its tool calls, the database commit and the notification that followed.

What is never logged

  • Card data, OTPs, tokens, passwords, secrets
  • Raw phone, email or address (hashed reference only)
  • Unredacted conversation text or full LLM prompts
  • Request bodies of payment and identity calls

Redaction

Logging libraries mask known sensitive fields by name. The pipeline then runs pattern detection for phone numbers, emails and card-like numbers, and CI tests fail if a sensitive field appears in log output.

Volume and cost control

  • Errors and warnings kept in full
  • Successful request logs sampled (for example 10%) with full traces kept for slow or failed requests
  • Per-service quotas with alerts on sudden growth
  • Tiered storage: hot, warm, archive

Access and integrity

Role-based access to log stores, with every query on conversation and audit logs itself recorded. Audit logs go to write-once storage with hash chaining so tampering is detectable.

Alerts from logs

  • Spike in INTERNAL_ERROR or a new error signature after a deploy
  • Repeated authorisation denials for one account
  • Any dead-letter or reconciliation mismatch entry
  • Redaction pipeline failures

Privacy requests

Because logs hold only a hashed customer reference, an erasure request is met by deleting the mapping key. Short retention covers the rest.

Tooling

OpenTelemetry for collection, a searchable store such as OpenSearch or Loki for hot logs, object storage for archive, and Grafana for unified views across logs, metrics and traces.

15

Deployment topology

flowchart TB
  U["Users: web, app, WhatsApp, voice"] --> CDN["CDN + WAF"]
  CDN --> GW["API gateway"]
  subgraph VPC["Cloud region, 3 availability zones"]
    GW --> BFF["BFF pods"]
    GW --> AG["NL agent pods"]
    BFF --> APP["Core app modules"]
    AG --> APP
    AG --> MG["Model gateway"]
    APP --> PG[("Postgres primary + replicas")]
    APP --> RD[("Redis cluster")]
    APP --> OS[("OpenSearch")]
    APP --> KF{{"Kafka"}}
    KF --> WK["Async workers"]
    WK --> PG
    SCH["Scheduler / workflows"] --> WK
  end
  MG --> LLM["LLM provider"]
  WK --> EXT["Payments, messaging, shipping"]
  KF --> DW[("Data warehouse")]
  

Figure 13. Stateless pods autoscale on CPU and queue depth; async workers scale independently from customer-facing traffic.

16

Key decisions and trade-offs

DecisionAlternativeWhy this wayCost accepted
Modular monolith firstMicroservices from day oneFaster delivery, simpler transactions, small teamShared deploy until split
LLM proposes, service commitsLLM writes directlyDeterministic validation, no invented slots or pricesMore tool plumbing
Redis hold plus DB constraintDatabase row locks onlyFast UX and a hard guaranteeTwo mechanisms to operate
Outbox plus event busDirect service-to-service callsLoose coupling, retries, replayEventual consistency for side effects
Precomputed availabilityCompute on every readRead path is 100x the write pathInvalidation on roster changes
On-device try-onServer-side renderingPrivacy and latencyQuality varies by device
Model gateway abstractionCall one provider directlySwap models, control cost, central guardrailsExtra hop
GraphQL BFFREST onlyEach channel fetches exactly what it needsSchema governance
17

Phased roadmap

Phase 1 · Foundation

Sell and book

  • Catalog, quiz, checkout
  • Guided booking form
  • Profile, notifications
  • Salon tablet basics
Phase 2 · Conversation

Book by chat

  • In-app NL booking agent
  • WhatsApp channel
  • Subscriptions and dunning
  • Membership and credits
Phase 3 · Intelligence

Anticipate

  • Voice booking
  • Virtual try-on
  • Proactive rebook nudges
  • Waitlist and smart overbooking

Success metrics

AreaMetric
BookingConversion from intent to confirmed, time to book, no-show rate, chair utilisation
NL agentContainment rate, turns per booking, fallback rate, wrong-booking rate (target near zero)
CommerceQuiz completion, shade return rate, subscription retention
ExperienceCSAT after visit, cross-journey adoption (home and salon)