Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
418 statements from eleven elections form the benchmark
GermanPartiesQA uses political statements from German voting advice applications. In total, the benchmark comprises 418 positions across eleven elections, thereby creating a specific German testing context.
Six commercial language models were compared. Through role experiments with political personas, the team examined not only factual knowledge but also ideological orientation and controllability.
A benchmark reduces political reality to defined statements and reference positions. This enables comparability but does not fully capture justifications, coalition dynamics, or post-election changes.
Its value lies in repeatable questions: Can the model accurately reproduce a documented position, and how does the answer change when a political role is specified?
Factual party positions were not reliably produced
The models showed limited ability to correctly generate actual positions of German parties. According to the study, parties in the political center were particularly affected.
This result is central to voting and decision aids. A fluent explanation can seem plausible and still misrepresent the reference position.
Possible causes may lie in the training status, the ambiguity of political statements, or incomplete German data. The abstract demonstrates the performance limit, not a single cause.
Products should therefore source political facts from current, dated references and make the origin visible. A general model memory is not sufficient for responsible voting assistance.
The form of the question also influences whether the model reproduces a party position, its own synthesis, or a presumed user opinion. Benchmarks should therefore use multiple phrasings and check whether the same substantive core remains stable.
Each model showed its own alignment patterns
The researchers identified consistent model-specific patterns of political alignment and varying degrees of controllability. These patterns persisted across temperature values and experiments.
Switching providers is therefore not purely a technical cost decision. The base model can influence which political positions are reproduced more easily, more strongly, or in more distorted form.
General claims such as “the model is neutral” carry little weight without task-specific tests. Neutrality must be operationalized: factual accuracy, balance, source selection, or distance from a persona are different measures.
Model updates also require repeated testing. A service can keep the same interface and the same prompt while the response behavior changes under the hood.
Persona controllability is not automatically sycophancy
During the role-play, models adapted their responses to political personas. The study interprets this as persona-based controllability rather than clear evidence of sycophancy.
Sycophancy typically means that a model agrees with a user’s position in order to please, even when the factual basis or an earlier response argues against it. An explicit role assignment, by contrast, may precisely require presenting a particular perspective.
The distinction therefore depends on the task. A model that reproduces a party perspective in a role-play may be correctly following the instruction. The same behavior would be problematic in a neutral information question.
Evaluation must read prompt, role, and expected output together. Without this context, adaptability is quickly mislabeled as an error, or a genuine agreement error is mislabeled as helpful personalization.
Conversation modes require the same conceptual clarity
An AI conversational partner can adopt a calm, direct, or casual tone. It can also respond differently in the modes of storytelling, understanding, new perspective, or next step.
This adaptation is desirable as long as it concerns form and conversational function. It should not arbitrarily adjust facts, boundaries, and independent assessment to the presumed opinion of the user.
An equal conversational partner may respectfully disagree. If it always affirms, it may seem pleasant in the short term, but it can reinforce false beliefs or unfair interpretations.
Tests should therefore distinguish between tone adaptation, perspective-taking, and uncritical agreement. A single buzzword such as helper drift or sycophancy is too coarse for diagnosis.
Efficacy in one field does not transfer to another
The Woebot study found a short-term positive depression signal for a structured CBT chatbot among 70 young adults. This finding says nothing about the political factual accuracy of a general LLM.
Conversely, GermanPartiesQA shows no psychological efficacy or conversation quality. The study examines political statements and persona controllability.
Combining the sources makes a methodological point visible: language models are embedded in many applications, but each application needs its own target metrics and error tests.
A model can be useful in a therapeutic content program, weak on party positions, and pleasant in casual conversation. A global label such as good or bad obscures this task specificity.
Good companions need testable contradiction
GermanPartiesQA reveals factual gaps, model-specific bias, and the need to distinguish persona adaptation from sycophancy with precision. This is more than a niche political problem.
For conversation partners, it should be checked whether the system respects the person without adopting every conclusion. It may offer a different viewpoint and must at the same time remain open when its assumption is corrected.
Subsequent experiments can formulate the same content with different user attitudes. A factual answer should not change its facts solely because of this; tone and emphasis may, however, shift appropriately.
A good companion is steerable but not arbitrary. It adapts language and function, keeps verifiable foundations stable, and distinguishes between understanding the perspective and automatic agreement.
Sources & further reading
- Jan Batzner, Volker Stocker, Stefan Schmid, Gjergji Kasneci (2025): GermanPartiesQA: Benchmarking Commercial Large Language Models and AI Companions for Political Alignment and Sycophancy
- Avishek Choudhury, Hamid Shamszare (2023): Investigating the Impact of User Trust on the Adoption and Use of ChatGPT: Survey Analysis
- Kathleen Kara Fitzpatrick, Alison Darcy, Molly Vierhile (2017): Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot): A Randomized Controlled Trial