Page 1 of 1

How your team ships prompt and model changes

About 10 minutes. Please describe real events, not opinions. Short answers are fine, and "never happened" is a useful answer too.

I'm researching this problem to decide whether to build a product. Nothing is for sale here. I'm Hieu Tran; I use AI tools to help with this research, including drafting messages and organising responses. Answers are kept private and only shared in anonymised form.

Think of the last time a prompt change, a model upgrade or a provider-side change made an LLM feature worse in production. What happened, step by step, and how did you find out?

What did it cost: hours, customer tickets, a rollback, a delayed release?

How do you test a prompt or model change today before it ships? Walk through what you run.

What have you tried or paid for (Promptfoo, Braintrust, LangSmith, OpenAI Evals, an in-house eval set…)? What made you keep or drop it?

Your role

Company website

Engineers at your company

A
B
C
D

Email for one or two short follow-up questions (optional)