Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The study distinguished three relationship logics
Tschopp and Sassenberg considered authority hierarchy, market-based exchange, and peer bonding. These categories describe whether the AI is experienced more as a serving tool, a calculating actor, or a socially close counterpart.
The cross-sectional survey targeted people with experience in voice shopping. Products with low and high personal involvement were distinguished.
For important purchases, socioemotional elements of peer bonding were more relevant. Calculative market relationships played a lesser role, while hierarchical master–servant perception was significant for less important purchases.
The results pertain to purchase intentions, not psychological support. However, they show that the task and the perceived relationship jointly influence how a voice assistant is used.
Relationship development correlates with trust
Seymour and Van Kleek surveyed 500 users of voice assistants. They examined whether human–device interactions can be described using Knapp's stage model of interpersonal relationships.
The data showed that this was possible and that more developed relationships were associated with higher trust and stronger anthropomorphization.
The correlation does not establish a clear direction. Trust can encourage more frequent and more personal use; conversely, repeated interaction can increase trust and anthropomorphization.
For long-term companions, this feedback loop is central. Every remembered name, every familiar voice, and every continuous session can reinforce the impression of a relationship.
Voice carries emotional form
The Emotional Chatting Machine models emotional categories, an internal emotional state, and an external emotional vocabulary for textual responses. In speech, prosody, volume, pauses, and tempo are added.
These features can make an identical sequence of words seem warm, sober, uncertain, or authoritative. Voice design therefore requires its own tests and cannot simply read aloud the finished text.
A synthetic voice should not overact emotion. Especially with vulnerable topics, strongly performed empathy can seem artificial, patronizing, or manipulative.
Options for tempo, voice, and interruption can give users control. The function must remain just as usable without audio.
Voice changes interruption and conversational rhythm
In text, the user usually decides for themselves when to send. In a spoken dialogue, the system must recognize whether a pause marks the end of the utterance or is merely a moment of reflection.
Responding too early can interrupt the user’s thought; responding too late feels sluggish. Especially when it comes to unburdening, the ability to say several sentences in a row is important.
An explicit function such as “I’m collecting my thoughts first” or a freely selectable listening mode can reduce the pressure of turn-taking. The system should not ask a new question after every brief pause.
Transcription introduces additional errors due to dialect, background noise, and emotional pronunciation. Correction must be simple and must not be interpreted as new psychological content.
For long utterances, a visible interim display can help without interrupting the flow of speech. The user should then be able to decide whether to review the transcribed text first or send it directly.
An accessible voice mode also needs clear feedback for the start, pause, and end of recording. These signals must not be acoustic only, so that people with hearing impairments or muted devices can retain control.
Voice data is particularly sensitive
Audio contains biometric and situational cues beyond the words themselves. Voice, surroundings, and third parties present can all emerge from a recording.
A voice chat must explain whether audio is processed solely for transcription, stored, or used for quality analysis. The text log and the raw recording are different types of data.
Local processing can reduce certain risks, but it is not available in every architecture. Data minimization means deleting raw data once its purpose has been fulfilled and no voluntary further use has been agreed upon.
Privacy is also needed on shared devices. An automatically read-aloud response can reveal sensitive content even though the technical data path is protected.
Voice is not an automatic quality gain
A voice can reinforce the feeling of being truly heard, but it does not repair a poor dialogue. A system that jumps to conclusions or forgets corrections does so just as much when spoken.
Voice adds latency, transcription, speech output, and turn-taking as new sources of error. Each layer requires monitoring and comprehensible failure behavior.
The comparison with text should therefore be conducted blind and task-specific. Does speech actually improve openness, satisfaction, or the flow of conversation, and for which people?
Some users prefer text because they formulate more slowly, need discretion, or feel less socially obligated without an audible voice.
Voice should therefore be a genuine option, not a premium tier of conversation. Switching between speaking and writing within the same session can better reflect different needs, provided the context is reliably preserved throughout.
Sources & further reading
- Marisa Tschopp, Kai Sassenberg (2024): The Impact of Human-AI Relationship Perception on Voice Shopping Intentions
- Hao Zhou, Minlie Huang, Tianyang Zhang (2018): Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory
- Seymour, W, Van Kleek, M (2021): Exploring interactions between trust, anthropomorphism, and relationship development in voice assistants