How model choice and response readability shape customer trust
IBM 6530, Cal Poly Pomona
2026-03-01
ODI’s chatbot has to solve a real customer experience problem, not just answer product questions.
Business need: Identify the chatbot configuration that feels clearest and most empathetic to customers.
This study tested whether ODI can improve chatbot service quality through:
The goal was not simply to make responses shorter. The goal was to identify the response style that feels clear, natural, and customer-first.
The framework turns chatbot design into a measurable business decision.
This gives ODI a scalable way to test service quality before live deployment.
The scoring rubric measured whether the chatbot:
Two LLM judges rated each response independently, with agreement checked using quadratically weighted kappa. This made the empathy score auditable and repeatable rather than subjective.
Two readability measures were tested:
The main regression tested whether readability predicted empathy while holding the rider question constant, with model and prompt included as controls.
Across configurations, model selection was the strongest driver of empathy.
Implication: ODI should treat model selection as the primary strategic lever.
FKGL showed a nonlinear relationship with empathy.
The strongest responses landed around a conversational high-school reading level.
Responses that are too simple can sound robotic and generic. Responses that are too complex can sound technical, verbose, and emotionally distant.
The best-performing responses were:
Moving from very simple language to the optimal range improved empathy by roughly 0.3 to 0.5 points on a 5-point scale.
Dale-Chall showed a weaker and less stable relationship with empathy.
Implication: simplifying vocabulary alone is not enough to improve ODI’s chatbot experience.
| Driver | Result | What it means for ODI |
|---|---|---|
| Model choice | Largest effect | Biggest lever for service quality |
| FKGL | Strong, nonlinear | Target conversational high-school readability |
| Prompt design | Smaller effect | Best used for refinement |
| Dale-Chall | Weak, not robust | Not a dependable standalone lever |
| R-squared | ~0.79 to 0.80 | The models explain most empathy variation in this dataset |
Validation included:
The core conclusion stayed consistent: readability matters, but model choice matters more.
The chatbot should be designed to:
The winning design is not the simplest response. It is the one that feels clear, relevant, and empathetic.
For ODI, the most effective chatbot strategy is to:
If ODI gets this right, the chatbot becomes a scalable brand touchpoint that improves confidence, conversion, and loyalty.
Friedman, D. B., and Hoffman-Goetz, L. (2006). A systematic review of readability and comprehension instruments used for print and web-based cancer information. Health Education & Behavior, 33(3), 352-373.
Hsu, C., and Lin, J. C. (2023). Understanding the user satisfaction and loyalty of customer service chatbots. Journal of Retailing and Consumer Services, 71, 103211. https://doi.org/10.1016/j.jretconser.2022.103211
Kull, A. J., Romero, M., and Monahan, L. (2021). How may I help you? Driving brand engagement through the warmth of an initial chatbot message. Journal of Business Research, 135, 840-850. https://doi.org/10.1016/j.jbusres.2021.03.005
Kumar, A., Poungpeth, N., Yang, D., et al. (2026). When large language models are reliable for judging empathic communication. Nature Machine Intelligence, 8, 173-185. https://doi.org/10.1038/s42256-025-01169-6
Parasuraman, A., Zeithaml, V. A., and Berry, L. L. (1988). SERVQUAL: A multiple-item scale for measuring consumer perceptions of service quality. Journal of Retailing, 64(1), 12-40.
Singh, S., Jamal, A., and Qureshi, F. (2024). Readability metrics in patient education: Where do we innovate? Clin Pract, 14(6), 2341-2349. https://doi.org/10.3390/clinpract14060183
Yun, J., and Park, J. (2022). The effects of chatbot service recovery with emotion words on customer satisfaction, repurchase intention, and positive word-of-mouth. Frontiers in Psychology, 13, 922503. https://doi.org/10.3389/fpsyg.2022.922503
ODI Project