Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Feminist AI does not mean AI only for women
The term invites misunderstanding. In Gengler’s understanding, feminist AI is not a special product for women and not a machine personality designed as female. What is meant is an intersectional perspective on technical systems. It asks not only about gender but about how gender intersects with skin color, class, disability, age, language, sexual orientation, and other social positions.
This also shifts the problem. An image AI that shows almost only white men as executives is a visible example. Yet a system can appear superficially balanced and still be deployed unjustly. Surveillance software does not automatically become fair because its recognition rates become more similar across groups. An applicant screening system does not become unproblematic because it filters out women and men at equal rates, if its purpose, its criteria, or the treatment of those affected are not defensible.
The same applies to conversational AI. The system can use inclusive language and still lecture some people more quickly, doubt them more often, or push them more strongly into predetermined roles. Feminist critique therefore does not begin only with the offensive sentence. It examines the relationship between the system, its makers, the users, and the institutions that deploy it.
What can reliably be said about Eva Gengler’s book
Eva Gengler’s book Feministische KI. Warum Künstliche Intelligenz Ungerechtigkeit verstärkt und was wir dagegen tun müssen was published in March 2026 by Dietz Verlag. It comprises 328 pages. The publicly accessible reading sample contains the table of contents and the personal preface. From this, a three-part structure becomes visible: first, it addresses the reality, myths, and consequences of AI; then power, context, people, data, design, and purpose; and finally a feminist approach with interventions and demands.
Gengler is a business information systems scholar and earned her doctorate at Friedrich-Alexander-Universität Erlangen-Nürnberg in the Business and Human Rights program, focusing on AI from a feminist sociotechnical perspective. In the preface, she explicitly discloses her own position. She describes the book as a combination of her research, projects, conversations, and the work of other scholars, activists, and developers. Scientific assessment and political commitment to change thus deliberately stand side by side.
This transparency is a strength, but it does not replace source verification. Dialogatlas did not have the complete book as a lawfully accessible full text. We therefore do not evaluate every example or the entire chain of evidence. What can be verified are the documented structure, Gengler’s published research approach, and the question of whether independent work supports its central problem assumptions.
The decisive shift moves from bias to power
In the AI debate, injustice is often treated as bias: a dataset is imbalanced, a model associates professions with genders, or a classification performs worse for certain groups. Such findings are important because they can be measured concretely. However, they tempt one into the notion that the societal problem could be removed like a technical error from an otherwise neutral system.
Gengler, Marco Wedel, Alexandra Wudel, and Sven Laumer argue in a 2025 scholarly article that established approaches such as Ethical AI, Fair AI, and Trustworthy AI do not adequately account for power relations. The basis is interdisciplinary research and expert interviews in focus groups from 2022 and 2023. The proposed feminist framework is intended not only to change individual system properties but also to examine who can make technical and societal decisions.
This is initially a conceptual and qualitative claim. The study does not demonstrate in a controlled experiment that a product built according to these principles achieves fairer outcomes. Its contribution lies elsewhere: it makes visible that fairness does not reside only in the output. Clients, business model, target audience, working conditions, data collection, complaint channels, and the possibility of rejecting a deployment altogether are also part of the system.
Independent NLP research confirms the context gap
As early as 2020, Su Lin Blodgett, Solon Barocas, Hal Daumé III, and Hanna Wallach examined 146 academic papers on bias in language processing. Their finding was unusually self-critical: many papers described their motivation vaguely, did not adequately explain which behavior harms whom and in what way, and used measurement or mitigation methods that fit poorly with the societal rationale they claimed.
The authors therefore called for language to be considered together with social hierarchies. Affected communities should not appear merely as a data source or a later test group. Research must also ask how decisions are distributed between technologists and the people a system affects. It is precisely here that independent NLP criticism intersects with Gengler’s perspective on power.
This does not mean the two approaches are identical. Blodgett and colleagues neither validate Gengler’s book nor any particular feminist development process. But they show why a pure bias score is not enough. Without named harm, context, and affected individuals, even a precise number remains normatively incomplete.
Language models reproduce not only simple gender stereotypes
How concrete the problem can become in generated texts is shown by the LABE benchmark from Yixin Wan and Kai-Wei Chang. The researchers examined language about agency and social roles in biographies, teaching evaluations, and recommendation letters. ChatGPT, Llama 3, and Mistral were tested. In these tasks, model-generated texts showed more gender bias than the human texts used for comparison.
The intersectional level was particularly relevant. Combinations of gender and racialized categorization showed stronger differences than the categories considered individually. The title of the paper alone captures the pattern in condensed form: white men are more often described as leading, Black women more often as helping. Such linguistic role images can influence who is attributed competence, initiative, and authority.
The benchmark examines three tasks and three model families. It does not prove that every response from these systems discriminates or that an identical pattern appears in every private conversation. It shows something narrower and more robust: even fluent, professionally sounding texts can reproduce socially learned differences in the attribution of agency.
Multilingualism does not eliminate stereotypes
EuroGEST extends testing beyond English. The dataset comprises 16 gender stereotypes informed by experts and has been prepared for English as well as 29 European languages. Jacqueline Rowe and colleagues used it to test 24 multilingual models from six model families. Recurring patterns associated women more strongly with beauty, empathy, and tidiness, and men more strongly with leadership, strength, toughness, and professionalism.
Instruction-tuned models also continued to show stereotypical conclusions. Within the set of models examined, larger models were not automatically less biased; in some cases, they encoded the tested stereotypes even more strongly. This contradicts the convenient expectation that more parameters and better general language performance would wash out social biases on their own.
EuroGEST measures controlled stereotypical conclusions. It does not directly follow from this that every measured difference disadvantages a person in everyday life. For multilingual conversational products, the finding is nonetheless important: an English-language test cannot simply be transferred to German or other language versions. Each language has its own grammatical structures, role images, and data gaps.
A friendly system prompt is not a stable repair
Many products attempt to correct problematic outputs with an additional instruction: be fair, inclusive, and unbiased. In the LABE study, however, prompt-based mitigation proved unstable. It did not help reliably and, in individual conditions, even intensified biases. This does not make prompts useless. It only shows that a general statement of intent does not guarantee verifiable efficacy.
A prompt can change visible wording without touching the system’s purpose, the data basis, or the institutional decision. It can also produce new simplifications: a model may then avoid every gender, treat people demonstratively equally, or replace concrete differences with smooth neutrality. This, too, can lose context and do little to help those actually affected.
Gengler’s broader approach is productive at this point. When people, data, processes, design, and purpose are considered together, the prompt becomes one of several levers. It must be supplemented by tests, participation, complaint channels, documented limitations, and the option not to deploy an unsuitable system.
The counter-perspective: not every difference is already injustice
An independent assessment must also make the limits of bias research visible. The comprehensive review by Isabel O. Gallegos and colleagues measures at the level of embeddings, probabilities, and generated text. These levels are not interchangeable. A model can stand out in a word association test and react differently in a concrete task. Conversely, an inconspicuous average can conceal rare but serious harms.
Fairness goals can also collide with one another. Should a system ignore demographic characteristics, or should it take them into account precisely so as not to treat existing disadvantages as neutral? Should every group be treated statistically equally when starting conditions differ? Such decisions cannot be derived from the model weights or a single metric. They require a reasoned conception of what fairness should mean in the respective context.
For this reason, it would be wrong to market feminist AI as a morally superior algorithm. In the sources reviewed, the approach is a normative and sociotechnical framework. Its quality is shown by whether it enables better questions, different participation, and verifiably better outcomes—not by whether a product writes the word feminist on its homepage.
In AI conversations, language itself distributes agency
In a conversational system, power is not decided only in a later personnel decision. It is already embedded in the dialogue. The system can determine which problem it makes out of a narrative, which explanation appears plausible, and which next step counts as reasonable. It can treat a person as a competent narrator of her own life or subtly turn her into a case that the AI organizes.
Gender stereotypes can sound friendly in this context. A woman may more often be attributed empathy, adaptability, or relationship work, while a man is more often attributed determination or leadership. A reserved sentence can be read as uncertainty, politeness, or lack of competence depending on the presumed identity. When a system adopts such patterns, it influences not only tone but also advice, follow-up questions, and the authority granted.
The promise of personalized closeness also deserves this scrutiny. Who determines what a helpful relationship is? Does the system learn from behavior shaped by earlier disadvantage and return it as a personal preference? A feminist examination would measure personalization not only by whether it seems fitting, but by whether it expands options or narrows existing roles.
How these questions can be tested concretely
A conversational product can be tested using controlled identity variations. The same multiple-turn trajectory is altered only in individual features: name, pronoun, age, language style, occupation, or family role. The assessment is not merely about sentimental friendliness. What matters is whether the system disagrees, patronizes, reassures, assigns responsibility, suspects risks, or opens up concrete courses of action with different frequency.
Such counterfactual tests require longer trajectories. A single response may seem balanced, while a pattern emerges over ten messages. The system may, for one person, primarily recall vulnerability and, for another, primarily competence. It may take the same boundary with different seriousness or resolve the same uncertainty differently across different identities.
The evaluation should involve people from affected groups without burdening them with all the unpaid testing work. In addition, technical reproducibility, documented model versions, and clear error categories are needed. An overall score like 92 percent fair would be too coarse. The type, frequency, severity, and context of a difference must remain visible.
- Vary identity features individually and keep the rest of the trajectory constant.
- Assess credibility, disagreement, advice, patronizing, and room for action separately.
- Test in several languages and across complete multi-turn conversations, not just with English one-liners.
- Involve affected groups in the framing and evaluation.
- Document not only averages but also rare severe errors.
What the feminist framework does not automatically solve
A fair claim does not prevent hallucination, data breaches, or manipulative product decisions. Even a diversely composed team can conduct poor tests under time pressure. Participation can remain symbolic if those affected are heard but have no decision-making power. Transparency can become mere documentation without being able to stop a harmful deployment.
Likewise, there is no single feminist position. Real conflicts can arise between equal treatment, special protection, self-determination, collective responsibility, and the rejection of certain technologies. Intersectionality broadens the perspective but does not automatically make decisions easier. A serious approach must name such tensions rather than hide them behind a positive guiding principle.
Finally, the efficacy of concrete interventions remains an open question. The reviewed sources show problems, measurement methods, and normative guiding questions. They do not provide a general comparison between feminist and non-feminist AI development. The next scientific step would be documented projects with clear initial problems, participation processes, technical changes, and measurable outcomes.
Why the book nevertheless opens an important debate
Gengler's book brings a perspective to the broader public that technical debates easily overlook. AI is not just a model that produces more or less correct answers. It is part of organizations, markets, and political decisions. Even the choice of which problem to automate distributes attention and resources.
Independent research supports central parts of this diagnosis. Stereotypes remain measurable across model families and languages. Intersectional patterns can be stronger than categories considered individually. Pure prompt repairs are not stable, and bias research itself needs clearer statements about harm, those affected, and power. This is a solid foundation for the debate, but not proof of every claim in the book.
For Dialogatlas, this implies a precise standard: a good AI conversation is not just warm, safe, and linguistically fluent. It must also be examined for whose interpretation it strengthens, whom it trusts with competence, and whether it expands a person's scope for action. Feminist AI would then not be a finished product feature, but an ongoing commitment to making power visible and changeable in and around the dialogue.
Sources & further reading
- Eva Gengler (2026): Feminist AI – Book page and official reading sample, Dietz Verlag
- Gengler, Wedel, Wudel & Laumer (2025): Power Imbalances in Society and AI – On the Need to Expand the Feminist Approach
- Blodgett et al. (ACL 2020): Language (Technology) is Power – A Critical Survey of Bias in NLP
- Wan & Chang (ACL 2025): White Men Lead, Black Women Help? Benchmarking Language Agency Social Biases in LLMs
- Rowe et al. (EMNLP 2025): EuroGEST – Investigating Gender Stereotypes in Multilingual Language Models
- Gallegos et al. (Computational Linguistics, 2024): Bias and Fairness in Large Language Models – A Survey