Pick a task with a cheap wrong answer
Drafting, categorising and summarising are ideal: a poor result costs seconds to correct. Anything where a wrong answer is expensive needs a human in the loop from the start.
Measure against the status quo
The comparison is not "is the AI good" but "is this better than what people do today". A draft that saves five minutes and is edited is a success; a perfect draft nobody trusts is not.
Budget for the boring parts
Evaluation, prompt versioning, rate limits, cost caps and a fallback when the provider is down. These take longer than the feature and determine whether it survives contact with real users.