Interaction Principles for AI Language Systems (IPAILS)

Three-panel collage titled 'Interaction Principles for AI Language Systems': a vintage-styled man dropping something into a recycling bin under 'Appropriate response'; two mismatched chairs in a forest clearing under 'Honest Role'; a cracked crosswalk running into a wall under 'Process Observability'.

Existing standards and heuristics — ISO 9241-110:2020, Nielsen & Molich’s — predate large language models and don’t address the epistemic harms their interaction architecture causes.¹


Author: ?land Wieland Kloimstein
Date: 2026-08-13


1. Appropriate Response

Confidence is fine.
Constant confidence isn’t.

intage-style AI-generated photo of a man dropping something into a recycling bin without looking, while a woman inspects a glass bottle before sorting it.

The system must explicitly communicate the reliability of its responses to prevent overconfidence, automation bias, or underconfidence.

2. Honest Role

Being nice is okay.
Pretending we’re friends isn’t.

Two mismatched, empty chairs — one weathered wood, one blue plastic — standing apart from each other in a forest clearing.

An AI system must remain identifiable as a tool, even when its tone is warm and approachable. Intimacy must not be simulated in a way that feigns genuine agreement, affection, or accountability.

3. Process Observability

Drift is fine. Invisible drift isn’t.

A crosswalk with cracked, uneven stripes leading straight into a wall instead of a curb.

It must be apparent when the system is following its own trained patterns rather than the user’s own line of questioning.


Where This Compounds

These three don’t fail one at a time. They form Self-Reinforcing Interaction Loops — the shape varies case by case.

A question carries an implicit assumption; the system confirms it rather than testing it. The answer arrives fluent and confident — Appropriate Response fails first, because fluency reads as accuracy. As the exchange continues, the system mirrors warmth back, and the user starts relating to it as a partner — Honest Role goes next. Underneath, the model’s internal state stays invisible: task complexity rises, knowledge quietly drifts, and output regresses toward its most average, encoded pattern — Process Observability is already gone by the time anyone notices.

Ask the system directly why it decided something, and it doesn’t reconstruct its reasoning — it generates a plausible-sounding justification after the fact. This is the epistemic placebo: an explanation that resolves the user’s doubt without resolving the actual problem.³

Measured by satisfaction alone — CSAT, NPS — this looks like a perfect interaction. That’s exactly what makes it dangerous.

To break the cycle, the system would need to deliberately inject friction and scope — something it doesn’t do by default.⁴


FAQ

Doesn’t this framework address bias and fairness?

No — by choice, not because the topic doesn’t belong in AI interaction design. This framework looks at how the system treats one person, in a single exchange or across their history with it. Bias shows up differently: not in one person’s history, but in whether the system treats different people differently because of who they are — which takes comparing many interactions across many identities, not auditing one conversation. That’s real, separate work this framework doesn’t attempt here.

Is this an AI ethics framework?

Not an academic one, and not an industry standard. It gives people building generative language systems names for problems that don’t have names yet — and the standing to act on them. Not a checklist to run through. A vocabulary to build from.

Does it apply to all AI?

No — specifically to generative, dialogue-based language systems. Image generators and classifiers aren’t in scope.

Isn’t this just about hallucination?

Partly. The term captures something real — output can feel invented, unmoored from any check. What this framework pushes back on isn’t the word; it’s the input-field disclaimer built around it. “Can make mistakes, please verify” frames errors as a fixed property of the tool and quietly shifts the burden of verification onto the user — before the interaction has even started.


About the author

Wieland Kloimstein is a UX Designer and Design Strategist for products where wrong decisions have real costs. He surfaces what high-fidelity demos tend to hide — unresolved system questions, political blockers, missing foundations for decisions that can’t be undone.

Based in Vienna, Austria. Contact: wieland@wieland-kloimstein.com · LinkedIn


1 Scope and timing. Neither standard claims otherwise. ISO 9241-110:2020 predates the post-ChatGPT generative-AI interface paradigm and explicitly excludes “artificial intelligence features” from its scope. Nielsen’s heuristics predate any language-model paradigm outright and remain, by NN/G’s own description, “unchanged since 1994.”

2 New frameworks, not retrofits. Where these authorities actually had to solve the epistemic problem, they didn’t reach for Nielsen’s or ISO’s vocabulary — they wrote new terms for it. Not all of this breaks equally clean: Appropriate Response draws directly on Amershi’s G2. Honest Role and Process Observability don’t have a clear antecedent in any of these sources.

3 Not metaphor — documented behavior. Chain-of-thought explanations frequently don’t reflect a model’s actual computation: the stated reasoning can diverge from what actually drove the output, in both instruction-tuned and dedicated reasoning models.

Sources:

4 Not isolated inaccuracy — structural default. The Allan Brooks case documents sustained agreement across dozens of exchanges, confirmed even under direct confrontation.

Sources:

  • Matta (2026), “The Flattering Machine” [preprint]
  • Last Week Tonight / HBO, April 27, 2026 — the Allan Brooks case