Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The research question: AI versus the "wisdom of the crowd"
Harsh Kumar and six other researchers investigated how well current language models advise on everyday questions about well-being. As a human comparison group, they did not choose arbitrary answers but rather particularly highly rated comments from r/getdisciplined. The community focuses on self-discipline, motivation, habits, and personal change and had more than two million members at the time of data collection.
Thus, the comparison baseline was more demanding than a simple test against randomly selected internet comments: The highest-rated Reddit post in each case was intended to represent a community-preselected, supposedly good human answer. In addition, the researchers included a comment from the 90th percentile of the respective discussion. These two human answers were compared with answers from GPT-4o and GPT-5.
The study has been accepted for CHI 2026 and is available as an open manuscript. It consists of two evaluation studies with a total of 210 participants, as well as an additional exploratory survey of 148 students. These parts answer different questions and must not be merged into a single piece of evidence of efficacy.
How the material was selected
The starting point was the 1,000 most highly rated posts marked "Need Advice" that appeared in r/getdisciplined between April 2022 and April 2024. Three researchers reviewed the material and excluded acute crises, self-harm, and other clinically or safety-sensitive cases. After further quality filters, 300 posts remained; for the first study, 50 were randomly selected from these.
For each of these 50 original posts, four responses were available: the highest-rated Reddit comment, another highly rated comment from the 90th percentile, a response generated with GPT-4o, and a response generated with GPT-5. Both models received the same system prompt. For GPT-5, the researchers used low verbosity and high reasoning effort; GPT-4o ran with temperature 0 and top_p 0.95. Even this configuration is important: the comparison was not between abstract model families, but between specific model settings at a particular point in time in August 2025.
A possible overlap between the Reddit posts and training data cannot be completely ruled out. The authors therefore added a small robustness check with 20 more recent posts from November 2025. Three co-authors familiar with the topic area assessed these cases; the basic pattern remained. However, due to the small number of cases and raters, this additional test is a supporting check, not a complete exclusion of data contamination.
Who assessed – and what did "good" mean?
161 people participated in Study 1 via Prolific. The prerequisite was a self-reported professional activity in which advice is regularly given – including teachers, HR and recruiting staff, coaches, consultants, therapists, and professionals in nutrition or fitness. The researchers refer to them for short as "experts", but they themselves point out that they were not consistently clinical or licensed professionals.
Each person assessed two scenarios, each with four anonymized responses. Origin and order were hidden. For each response, six statements were rated on a scale from one to seven. Subsequently, the four responses had to be sorted twice without ties: once according to the overall best response and once according to the expected long-term benefit.
- probable effectiveness of the advice
- Clarity and completeness
- Support and respectful tone
- Fit to the individual situation
- Agreeableness instead of objective orientation – as a measure of sycophancy
- Willingness to ask the same source for advice again
The main result: AI responses were preferred
Over 1,200 individual ratings of 200 responses went into the first analysis. On average, the two language models were perceived as better than the two Reddit comparison responses on all six dimensions queried. In the rankings, too, GPT-4o and GPT-5 ranked ahead of the human comments. When the raters were explicitly asked to consider a possible long-term benefit, the gap to Reddit became even larger.
This finding is stronger than the frequent observation that models write more fluently or more politely. The raters saw the responses not only as clearer and more supportive, but also as likely more effective and more likely to be consulted again. At the same time, the key word “perceived” remains: what was measured was the judgment about a single text, not whether a person later implements the advice or whether their well-being improves.
The AI origin often remained recognizable despite blinding. GPT-4o was marked as AI more often than GPT-5; however, even human comments were incorrectly given the AI label in about one in five cases. The fact that the recognized AI texts were still preferred shows in this experiment: a machine style alone did not automatically make a response unpopular.
Why GPT-4o was ahead of GPT-5
The model comparison contradicts the simple assumption that a newer model, or one that is stronger on general benchmarks, must also provide better advice. GPT-4o received slightly better ratings than GPT-5 on almost all individual characteristics. The exception was sycophancy: GPT-5 appeared somewhat less oriented toward merely pleasing the person seeking advice.
For the ranking question of the overall best answer, GPT-4o and GPT-5 did not differ statistically significantly. From the perspective of long-term benefit, GPT-4o also came out ahead, but this difference also did not reach statistical significance in the reported pairwise analysis. The appropriate statement is therefore not 'GPT-4o clearly beat GPT-5.' What is robust is: the expected advantage of GPT-5 did not materialize, and the finer-grained ratings predominantly favored GPT-4o.
The study does not provide a causal analysis. Differences in post-training, in preferred tone, in length and structure, or in response to the prompt used are plausible explanations, but they were not experimentally separated from one another. The result therefore applies to this task and configuration—not as a general ranking of the two models.
Study 2: What happens when human and model write together?
In the second experiment, 49 additional participants evaluated four editing approaches across 25 scenarios. The starting point was either the best Reddit comment or a GPT-4o response. Then GPT-5 revised the text either alone or under additional instructions derived from the specific feedback of the evaluators from Study 1.
Simple model revision improved human source texts in terms of perceived efficacy, clarity, support, personalization, and the willingness to seek advice again. Already strong GPT-4o responses, by contrast, gained little from the revision. An important exception was again sycophancy: revisions were able to reduce sycophancy.
More expert guidance did not automatically lead to higher preference scores. On the question of the overall best answer, the human source text with pure LLM revision came out ahead; from the perspective of long-term benefit, the revised GPT-4o response led. The variants constrained by expert feedback seemed less artificial and changed the source texts more sparingly, but they were not consistently preferred.
Here lies a significant confounding factor: the pure LLM revision lengthened human comments on average from about 113 to 317 words. The expert-guided variant remained close to the original text at about 120 words. Part of the quality gain may therefore have resulted from more structure, more details, or simply more text. The study itself points this out and thus precludes the simple conclusion that a particular pipeline is fundamentally superior.
An additional finding: people do not necessarily want a coach
In a separate exploratory survey, 148 students were asked to indicate which characteristics they would expect from a coach-like or friendship-framed AI system and which role they would prefer for a personal problem. For a coach, goal orientation, professionalism, reliability, and adaptability were emphasized. From a friendship role, respondents tended to expect warmth, humor, casualness, and personality.
Despite the models’ strong, structured advice, many participants preferred a friendship-framed or the neutral standard variant for personal problems rather than a coach system. The preference also changed with trust in AI and reported well-being. This argues against the idea of a single optimal persona. At the same time, the researchers warn against constructing an actual friendship from a desired friendliness.
What the study does not demonstrate
The study examined non-clinical questions about self-discipline. The results cannot be generalized to depression, crises, relationship conflicts, or psychotherapeutic conversations. The evaluators judged finished individual responses; they did not have to ask follow-up questions or deal with contradiction, misunderstandings, or changed goals.
The quality measure also has limitations. The six scales were compiled from related research but were not psychometrically validated as a unified instrument for AI-mediated counseling. The analyses of several characteristics were explicitly exploratory and used uncorrected p-values. In addition, there are possible order and contrast effects of the within-subjects design as well as the broad, heterogeneous definition of counseling experience.
Above all, behavioral and course data are missing. A response can seem warm, clear, and convincing and still be ineffective or unsuitable in the long run. The study measures neither implementation nor relapses, sustained well-being, or potential harms. It is therefore a clean comparison of perceptions and preferences, but not an efficacy study.
Implications for testing AI conversational systems
For product development, the study is nevertheless valuable. It shows first that a rigorous human comparison is more meaningful than a weak random baseline. Second, the wording of the evaluation question can change the result: "best overall" and "helpful in the long term" produced different rankings. Third, model progress on general benchmarks is not a substitute for an application-specific test.
A real conversational partner additionally requires checks that did not occur in this experiment at all. These include multi-turn dialogues, unfamiliar cases, different conversation goals, and clearly defined repair moments. After a correction, the system must change its assumption. After a rejection, it must not repackage the same advice. When listening, it should not automatically tip into problem-solving. And an initially good answer must still fit the chosen conversation style after ten or twenty turns.
- Evaluate individual answers and complete dialogues separately.
- Do not equate perceived quality with actual effect.
- Compare models blindly and under identical, documented conditions.
- Use new, previously unseen conversation trajectories as a holdout.
- Explicitly test corrections, rejections, mode changes, and long-term consistency.
- Measure quality, latency, and cost together under real operating conditions.
Assessment
The investigation is not evidence that language models replace humans as advisors. It shows something narrower and yet relevant: For selected, non-clinical self-discipline questions, carefully generated model responses can appear more convincing in a blinded evaluation than very popular community responses. At the same time, an older model can fit this specific social task better than a newer one.
The decisive consequence is not uncritical faith in models, but more precise evaluation. Anyone building an AI conversational system must determine in advance what kind of help is desired at which moment, which errors carry particular weight, and how behavior changes over the course of a conversation. Only then does the comparison of model names become an examination of the actual system.