Explanation: AI Integration & Tenant Security
Architectural explanation of the Vercel AI Gateway tool-calling agent, context window packing, and tenant isolation boundaries
The AI Assistant provides conversational guidance, standard operating procedure assistance, and drafted communications for restaurant operations. It utilizes DeepSeek V4.1 Flash through the Vercel AI Gateway (@ai-sdk/gateway) with the Vercel AI SDK and structured system prompts. This document details the architectural boundaries, tenant isolation, and prompt context strategy.
Architecture Overview
The AI Assistant has no configured Mastra memory. Conversation history and generation state are managed by Danvas. The app's /api/assistant/chat BFF forwards authenticated context to the API's canonical /api/v1/assistant/chat/v2 runtime, which revalidates scope and builds role/mode instructions.
Execution runs through the canonical Mastra runtime
(apps/api/lib/assistant/mastra.ts). There is no current runtime gate named
FEATURE_MASTRA_ASSISTANT; bounded features have their own admission and feature
gates. See the workload routing reference.
graph TD
A[User Chat Message] --> B[Next.js /api/assistant/chat Handler]
B --> C[Extract getActionContext: teamId, userId, role, locationIds]
C --> D[Authorize request context: teamId, userId, role]
D --> E[Build Mode & Context System Prompt]
E --> F[Invoke Vercel AI Gateway model stream via Vercel AI SDK]
F --> G[Stream SSE text chunks to Client]
G --> H[Persist Thread, Message Parts & Generation Telemetry]Tenant Security & Safety Boundaries
- Authentication Gate: The
/api/assistant/chatroute is protected bygetActionContext(), extracting the verifiedteamId,userId, androle. - Contextual Isolation: Server-side authorization and tool/source reader checks enforce role and location boundaries; prompt instructions explain those limits to the model. An admin with no explicit location assignments has team-wide location scope; scoped admins, managers, and members remain limited to their authorized locations. The native assistant route also revalidates forwarded focus locations against the team location map.
- Draft-First Mutation Boundary: Operational writes (announcements, incident filings, reports) are strictly draft-first. The AI produces editable structured text that users review and publish through native Danvas workflows; it never executes unverified database mutations.
- PII Redaction: Observability log scrubbing (
@repo/observability) appliesredactPIIfrom@repo/security/piito persisted logs, and prompts minimize staff and guest PII in generation context. - Abuse Control: Chat requests are rate-limited to 20 requests per minute per user via the Upstash Redis limiter (
packages/rate-limit).
Telemetry & Cost Accounting
Token usage and estimated inference costs are recorded in the chat_generations table:
- Provider-reported usage is settled from the Mastra stream with a bounded wait; missing counts remain unknown.
- Estimated costs are derived from the serving model's reviewed rate card; unknown usage is not recorded as zero cost. This chat coverage does not establish coverage for every direct generator or external provider.
Knowledge Retrieval Pipeline
Answers grounded in operational knowledge run through a hybrid retrieval pipeline
(createDatabaseKnowledgeSearch in apps/api/lib/assistant/database-search.ts):
- Database keyword and vector candidates are fused with reciprocal-rank fusion.
- When
KNOWLEDGE_RERANKING_ENABLEDis enabled, at most 20 bounded candidates are re-scored withvoyageai/rerank-3through OpenRouter (AI SDKrerank()via@repo/ai). The gate defaults off. - Reranking is fail-open: provider errors, missing keys, or cancellation return the fusion order unchanged, so search degrades instead of failing.
Embeddings use voyageai/voyage-4 through OpenRouter. Retrieval is bounded to the requesting user's team and
location authorization.