How to Test AI-Generated Responses
There are several important areas to validate.
1. Accuracy
The first question should be:
Is the information correct?
Suppose the application’s return policy is:
Products can be returned within 7 days.
The user asks:
“Can I return my product after 10 days?”
AI responds:
“Yes, you can return the product within 30 days.”
This is incorrect.
QA Result
❌ Fail
The AI response does not match the actual business rule.
2. Relevance
The response should answer the user’s actual question.
User:
“How can I reset my password?”
Good response:
“Go to the login page and click ‘Forgot Password’. Enter your registered email address and follow the instructions.”
This is relevant.
Bad response:
“You can update your profile information from Settings.”
The information may be related to the application, but it does not answer the user’s question.
QA Result
❌ Fail
3. Completeness
Sometimes an AI response may be partially correct but miss important information.
For example:
User:
“How do I return my order?”
AI response:
“Open My Orders and select Return.”
But according to the application requirements, the user must also:
- Select the product.
- Choose a return reason.
- Upload an image if required.
- Submit the return request.
The AI response is incomplete.
QA Result
⚠️ Partial / Fail depending on the requirement.
4. Consistency
AI responses may change between requests.
Ask the same question multiple times:
“What is the return period?”
Response 1:
“You can return the product within 7 days.”
Response 2:
“The return period is 10 days.”
Response 3:
“Products can be returned within 7 days.”
Now we have a problem.
The AI is providing different information for the same business rule.
QA Result
❌ Fail
For business-critical information, the response should remain consistent with the approved source.