System prompts define role, boundaries, and conversational style before a human writes the first word. New research shows what transparency and control users want—and why a published prompt alone does not yet yield a reliable AI.
A system selected suitable responses from real peer-support data. 79.2 percent were deemed acceptable, yet people still preferred responses framed as human—even when the text was identical.
STEF connects earlier emotional changes with the development of support strategies. On ESConv, the approach outperformed comparison models—a benchmark success, not clinical evidence of efficacy.
MemoryBank stores, reinforces, and forgets information from earlier conversations. This improves personal continuity but makes selection, correction, privacy, and false memories central product issues.
For experienced voice shoppers, peer bonding was particularly relevant for important purchases; a survey of 500 users linked relationship development with trust and anthropomorphization. Voice is more than just an input channel.
A survey of 500 users was able to describe human-device interactions using a relationship stage model. Advanced relationships were associated with greater trust and anthropomorphization.
Knowledge-grounded dialogue models with retriever, ranker, and generator significantly reduced hallucinations. For real-world chats, part of the problem thus shifts to search, source quality, and context assembly.
Language models can evaluate responses quickly and according to fixed criteria. Three studies also show why an automated judgment neither replaces human assessment nor an application-specific dialogue test bench.