We launched an AI feature to 100,000 users. Here's what went wrong and what we'd do differently.
What went wrong:
The LLM cost spiked 4x our projection because users sent much longer inputs than our test cases. Some users discovered they could make the system generate unlimited content by phrasing requests cleverly.
Edge cases we never tested: users who input entirely in emoji, users who pasted entire PDFs as text, users who tried to jailbreak the system on day one (took 23 minutes for someone to try).
The feedback pipeline was overwhelmed. We got 3,000 pieces of feedback in the first week and had no automated way to categorize them.
What we'd do differently:
Soft launch to 1% of users first. Every percentage point of traffic reveals new failure modes.
Set hard limits on input length, output length, and requests per user per hour BEFORE launch.
Build automated feedback categorization from the start, not after you're drowning.
Have a "kill switch" — the ability to disable the feature in 30 seconds if something goes catastrophically wrong.
And staff customer support specifically for AI-related questions. Users ask things about AI features that existing support scripts don't cover.
Every launch teaches you things testing never will. The key is making those lessons cheap by starting small.