← All learning notes

AI practice

Create a small evaluation set before changing a prompt

By SI100x · Published

Without stable examples, a prompt change can feel better while making other cases worse. A small evaluation set gives you a repeatable way to compare the behavior you care about.

Try this exercise

  1. Collect invented examples covering normal input, missing information and a confusing case. Write the expected qualities of a response for each one.
  2. Run the current prompt and save the outputs with the prompt version. Evaluate factual consistency, missing details and whether the requested format is usable.
  3. Change one part of the prompt and repeat the same examples. Add a new case when you discover a failure, rather than replacing the old cases with easier ones.

Check your result

A useful comparison shows both improvements and regressions. The set is evidence about those examples, not proof that the system is reliable on every input or safe for every use.

Explore your next step

Compare course curricula or try a free live demo to ask about prerequisites and practice work.

Explore courses →Find a free live demo →