Recent research describes a recurring problem in personal AI conversations: models sound empathetic but fall back into the same supportive role across very different situations. The missing quality is not more warmth, but a conversational posture that fits the moment.
Eva Gengler’s book shifts the AI debate from isolated biased outputs to power, purpose, and participation. Independent research supports this critique—but not the notion that feminist AI is already a finished technical method.
A new preprint separates helpful regulation from unwanted escalation in 9,000 model responses. During venting, both increased simultaneously: a response could appear attentive and still reinforce certainty, partisanship, or emotional intensity.
A mental-health chatbot can seem cautious in its first response and still overstep its role later. Two studies show why conversational boundaries must be examined across entire conversation trajectories.
A new experimental study shows: Even the mere availability of mostly incorrect AI hints can deter people from admitting uncertainty. What this means for personal conversational systems—and what the study does not prove.
A systematic review of studies on medical chatbot advice reveals massive reporting gaps: the model version was almost never clearly stated, prompt development was rarely described, and success was frequently assessed subjectively.
Specialist physicians assessed ChatGPT’s responses to nine questions about nasal surgery as understandable and informative. The study shows potential for initial information—but no personalization, no open dialogue, and no robust clinical advice.
ChatGPT answered nine questions about venomous snakebites in a way that was understandable and professionally convincing. But precisely the potential use in remote regions reveals the limitation: region, timeliness, individualization, and the risk of delay are all part of answer quality.
ChatGPT answered common questions about antiretroviral therapy correctly and identified a potentially life-threatening hypersensitivity reaction. As a supplement, a chat can lower barriers to access—but it does not replace individual, regional, or pregnancy-related counseling.
CORTEX classifies English and Polish texts by mood and nine emotions. The good classification scores demonstrate technical potential—and at the same time, why a label such as “sadness” does not yet determine an appropriate conversational response.
Using the example of retirement planning, Lo and Ross demonstrate three general hurdles for language models: domain-specific adaptation, trustworthiness, and regulation. The questions extend far beyond financial advice.
In an analysis, GPT-based chats were perceived as conversationally strong but less empathetic than humans. The apparent paradox shows: empathy does not arise from individual formulas, but from alignment over the course of the conversation.
In an experiment on HPV vaccination communication, the chatbot was overall similarly useful to a human. However, satisfaction declined with anger, and under shame, participants opened up more to humans.
In an interactive experiment, people followed AI recommendations even when context and their own judgment argued against doing so. Overconfidence harmed not only themselves but in some cases also third parties.
Conversational Automated Program Repair combines new code suggestions with test results from previous attempts. The technical principle is also instructive for chats: feedback must change behavior, not just be mentioned.
Two experiments show: sympathy and empathetic phrasing were preferred over a purely factual chatbot in sensitive health advice. What matters most is which kind of empathy fits the situation.
Empath.ai combines text, voice, and facial recognition with CBT content. The concept promises more context but introduces new uncertainties, privacy concerns, and requirements for evidence.