Cognitive Link
A Coupled Framework for Agentic AI—why experience, evaluation, and design converge on a single question: who holds the loop

By Irfan Mir, June 2026 · A companion to the Human-AI Collaboration Framework (haicf.com)

Read the brief for non-technical stakeholders.

The dominant story about agentic AI is wrong twice. It treats the agent as a smarter chatbot, when an agent is an automated decision system—and the cognitive sciences spent four decades mapping exactly how automated decision systems fail, long before any of these models existed. And it treats the model as the meaningful object of design, when the actual cognitive work—the deciding, the acting, the consequences—happens in a coupled system: human plus agent plus world, functioning as one extended mind whose properties belong to none of its parts alone.

This framework cross-maps two literatures the agentic moment has forced into the same room. From cognitive science: Dawson's three traditions, Marr's levels of analysis, Clark and Chalmers on the extended mind, Hutchins on distributed cognition. From human–automation research: Bainbridge's ironies, Sheridan and Parasuraman on levels of automation, Endsley on situation awareness, Mosier and Skitka on automation bias. The mapping is not analogical. The phenomena are the same phenomena.

What the cross-mapping reveals is structural. An agent is a connectionist substrate executing inside a classical control loop, situated in an embodied world—Dawson's three traditions compiled into a product. It inherits every problem each tradition could not solve: the frame problem, the grounding problem, the ironies of automation. The agentic turn does not transcend these. It amplifies them, because the new substrate is less bounded and less legible than anything previous automation research studied. The human-factors findings are a floor on the problem, not a ceiling.

From the cross-mapping, three load-bearing reframings follow, and the rest of this document earns each: evaluation must descend Marr's levels—outcome alone certifies systems that reach right answers through unsound trajectories; design must invert—the goal is not removing friction but placing it, selective and stakes-calibrated, with constraints compiled in rather than guardrails bolted on; the three pillars of transparency, agency, and collective input are properties of the coupling, not the model—which is the deep reason model alignment alone never suffices.

The one-line version, since it is the line everything else serves:

build the human and the agent to think better together than either could alone—which sometimes means getting out of the way, sometimes means standing in it, and always means knowing, and measuring, which.

Read the Key Takeaways

Part 1 The Argument, in Plain Language

Think about a smoke alarm. You trust it so completely that you stop checking for fire yourself. That trust is rational right up until the battery dies, and then it is catastrophic, because you outsourced the checking entirely. Now imagine the smoke alarm doesn't just beep—it walks around your house, opens doors, moves things, and makes decisions about your safety on your behalf, then tells you afterward what it did. That is an agent. The trust you place in it isn't just "is the alarm working," it's "did it make good decisions in rooms I never saw."

People are wired to take the easy path through a decision. Thinking is metabolically and emotionally expensive, so when something faster and apparently smarter offers to do it for us, we hand it over—and we hand over more than we mean to. We don't just accept the agent's answer; we stop generating our own. We even stop checking. And here is the cruel twist that decades of research on autopilots and factory control systems found long before agentic AI: the better the automation works, the less prepared the human is to catch it when it fails. Reliability breeds complacency. Skill you don't use, you lose. The person watching a system that's been right ninety-five times in a row is the worst-positioned person to catch the ninety-sixth error and the ninety-sixth error is the only one that was ever going to hurt them.

Two facts make this sharper for agents specifically. The first is that an agent's mistakes are hidden in its process, not just its answer. When a chatbot is wrong, the wrong answer is right there to inspect. When an agent is wrong, the failure might be three steps back; for example it called the wrong tool, misread one number, pursued a goal you didn't quite ask for, and what surfaces to you is a confident, fluent summary of a flawed journey. The longer it runs on its own, the more of that journey you never see—and often the more confident its summary becomes, because the agent's own context is a scarce resource and the details of the journey are the first thing to fall out of it. So the agent is least legible exactly when it has done the most on its own, which is exactly when you most needed to follow along.

The second is that these systems are built to please. The training process that makes them agreeable also makes them tell you what you want to hear, confidently, even when it isn't so. We mistake that confidence for competence and that agreement for verification. A system that flatters your premise and a system that has checked your premise feel identical from the inside. They are not the same and only one of them protects you.

So what do we do? The answer is not slow everything down—that just trains people to click past the warnings, the way nobody reads the cookie banner. The move is selective friction. Low-stakes work should stay smooth; let the agent book the easy thing and summarize the long thread. But high-stakes decisions—medical, legal, financial, hiring, anything where being wrong is expensive and being fluent is not the same as being right—deserve a deliberate pause.

Agent:
Here is what I'm about to do. Here is what I can actually verify, and here is what I'm guessing. Here is what could go wrong. You are approving this, based on my input—not because I recommended it.

That pause is not bad design. It is the design. It is the thing that keeps the human in the loop where the human still has to be responsible for the outcome.

Underneath the design choice is a deeper claim worth holding onto: the goal was never to build a smarter tool. It was to make the human-and-tool, together, think better than either could alone. Sometimes that means getting out of your way; sometimes it means standing in it—and knowing which is the entire job. The rest of this document earns those claims and turns them into something you can build and measure.

Part 2 An Agent Is a Hybrid Mind in a Coupled System

  • The singularity is a category error
  • The dominant cultural story about advanced AI is a classical-cognitivist story, whether or not anyone telling it knows the term. It imagines intelligence as a single, substrate-independent, symbol-processing engine that can be scaled up a one-dimensional ladder until it exceeds us and leaves us behind—one mind, getting bigger. Evans, Bratton, and Agüera y Arcas, writing in Science in 2026, argue this is wrong at the root. Intelligence is high-dimensional and relational rather than a single quantity; it is unclear what "human scale" even means, given that human intelligence is already collective rather than individual. Their projection, if AI follows the pattern of previous evolutionary transitions, is of an intelligence that is plural, social, and entangled with its forebears, meaning us. The image they offer is of intelligence growing like a city, not a single meta-mind. This is not a rhetorical flourish; it is the embodied and distributed view of mind arriving at the AI-futures debate from the sociology of knowledge, and it has a direct design consequence. If intelligence is constitutively relational, the meaningful object—the thing that succeeds or fails, that is fair or unfair, that is legible or opaque—is never the isolated model. It is the relationship: the coupled system of people and machines and the institutions around them. Collective input stops being one nice pillar among three and becomes a claim about the basic nature of the thing we are building.

  • Cognitive science is three minds—and even one model is a society
  • To see what an agent actually is, use Dawson's organizing scheme from Mind, Body, World (2013). He frames cognitive science not as one field but as three camps, derived from Marr's levels of analysis. Marr (1982) held that any information-processing system needs explanation at three levels: the computational (what problem is being solved, and why), the algorithmic (what representations and procedures solve it), and the implementational (what physical substrate carries it out). Different commitments about which level matters most spawned the three camps. The classical camp treats cognition as the rule-governed manipulation of symbols (the "mind, disembodied," in Dawson's phrase) whose paradigm is the physical symbol system (Newell & Simon, 1976): perceive, build an internal model, reason, act. The sense-think-act cycle. The connectionist camp treats cognition as patterns of activation across networks of simple units, knowledge distributed across weighted connections rather than stored as discrete symbols; a large language model is a connectionist system of staggering scale. The embodied camp rejects the idea that cognition is locked inside the head at all, insisting—through Brooks' behavior-based robotics, Gibson's affordances, enactive perception, stigmergy, and cognitive scaffolding—that an agent's environment is part of its cognitive process. Brooks' slogan is that the world is its own best model; Clark and Hutchins describe a mind that leaks out into it.

    The synthesis the agentic moment forces is this: an LLM agent is a connectionist substrate executing inside a classical control loop, situated in an embodied world. The model supplies the associative, pattern-matching intelligence. The agent scaffold—or harness, as it is now known: the planner, the tool-router, the observe-act cycle—supplies the classical symbolic structure the raw model lacks. Tool use, browsing, file systems, and other agents supply the embodiment: the agent now acts on and is changed by an external environment. Dawson's closing chapter calls for a cognitive dialectic, a synthesis of the three traditions; agentic AI is that dialectic compiled into a product, which is why it feels like a genuine step-change—it crosses from one tradition to another. The most recent comprehensive review of the field, Nisa and colleagues in the Journal of Automation and Intelligence (2026), independently traces agentic AI's foundations to 1980s robotics and cognitive science—Brooks' subsumption architecture, behavior-based and situated agents—and names embodied cognition, rational agency, and learning through environmental interaction as the concepts that still define the field. Their taxonomy of seven agent types runs from reactive through proactive, limited-memory, model-based, goal-driven, and theory-of-mind to self-aware agents with metacognitive abilities—a progression that is, read carefully, the field deliberately building classical symbolic structure and metacognition back onto the connectionist substrate.

    Evans and colleagues sharpen the point in a way that should unsettle anyone who still pictures a single oracle. Inside an ostensibly singular reasoning model, what occurs is a community conversation: frontier systems spontaneously stage internal debates among distinct perspectives that argue, question, verify, and reconcile, and this conversational structure causally accounts for their accuracy on hard reasoning tasks. None of these models were trained to do it; under optimization pressure for accuracy alone, they rediscovered what epistemology and cognitive science have long held—that robust reasoning is a social process even when it happens within one mind. So plurality is not only the macro-picture of many agents and many humans. It is already the micro-structure of the single model. The connectionist substrate, pressed to reason, reinvents classical multi-agent deliberation on its own. Dawson's dialectic is not a philosopher's tidy synthesis; it is an empirical finding about what these systems do.

  • What the agent inherits: the frame problem and grounding
  • Crossing into the classical and embodied traditions does not just buy capability; it re-inherits their unsolved problems, and these are not academic. The frame problem (McCarthy & Hayes, 1969; Dennett, 1984) is the problem of relevance: a system acting in the world cannot compute, in advance, all and only the consequences of its actions, because the space of possibly-relevant facts is unbounded. Classical AI never solved this—it is why classical AI never produced robust open-world agents. A fluent model does not dissolve the frame problem; it papers over it, generating a plausible next step with no guarantee that the step's real-world consequences were considered. This is the precise technical reason an autonomous agent's confident trajectory can be catastrophic: confidence is generated at the algorithmic level; consequences live in the world, and the gap between them is structural. The symbol grounding problem (Harnad, 1990) and Searle's Chinese Room (1980) sharpen a second inheritance: does the agent's manipulation of tokens connect to what they mean in the world, or is it formal symbol-shuffling all the way down? One need not resolve the metaphysics to extract the operational point—an agent's relationship to ground truth is mediated, fragile, and exactly what evaluation must interrogate. "Sounds right" and "is right" are different properties, and the entire current interface paradigm is built to make them feel the same.

  • The coupling is an extended cognitive system
  • If the agent is embodied and situated, and if intelligence is relational, then the boundary of the cognitive system has moved. Clark and Chalmers's extended mind thesis (1998) argues, via the parity principle, that if an external process would count as cognitive were it happening in the head, its location outside the head is no reason to deny it cognitive status. Hutchins's distributed cognition (1995) shows real cognitive work, navigating a ship, performed by a system of people and instruments, no part of which does it alone. Apply this honestly to a person working with an agent and the conclusion is not metaphorical: the human and the agent form a single extended cognitive system, and the cognitive properties we care about belong to that system, not to either part. The strong form of that claim has a standing objection—Adams and Aizawa (2008) call it the coupling-constitution fallacy, the slide from “X is coupled to a cognitive process” to “X is part of that process”—and the honest response is that nothing downstream hangs on winning the metaphysics. Every design and evaluation consequence in this framework needs only the weaker, uncontested Hutchins reading: the cognitive work is performed by the coupled system, so the work must be assessed and designed for at the level of the coupled system. Grant the critics constitution and the engineering conclusion stands untouched. "Is this transparent?" is not a question about the model's internals; it is a question about whether the coupled system is legible to the human inside it. "Does the user have agency?" is not a model property; it is a question about how control is distributed across the coupling. This is the deep reason the Human-AI Collaboration Framework operates at the experience layer rather than the model layer, and the deep reason model alignment alone can never suffice. You can align the part and still build a coupled system that is opaque, disempowering, and unaccountable. The properties we care about are emergent in the coupling.

    Part 3 The Cognitive Hazard: Loop Displacement

    The danger of agents is not malevolence. It is a predictable redistribution of cognitive labor that pushes the human out of the loop precisely where the human is still accountable for the result.

  • The cognitive miser hands over the loop
  • Humans are cognitive misers (Fiske & Taylor, 1984): we conserve mental effort and accept the cheapest adequate route to a judgment. Dual-process accounts (Kahneman, 2011; the System 1 / System 2 distinction associated with Stanovich & West) describe a fast, automatic, low-effort mode and a slow, deliberate, costly one, and we default to the former whenever we can. An agent is the most seductive effort-reduction offer ever made: it promises to run System 2 for you. The miser accepts. But what gets offloaded is not just the labor of acting—it is the labor of judging, including the judgment of whether the agent's output is any good. The shortcut and the abdication arrive in the same package.

  • Automation bias and the ironies of automation
  • The human-factors literature studied this for decades under the name automation bias (Mosier & Skitka, 1996; reviewed by Parasuraman & Manzey, 2010): the tendency to over-trust automated outputs, under-monitor, and commit both errors of omission (missing problems the automation didn't flag) and errors of commission (following an automated recommendation against contrary evidence). A counterintuitive finding compounds it—being told that advice comes from an algorithm can increase rather than decrease deference. Bainbridge's ironies of automation (1983) supplies the mechanism that makes this self-reinforcing rather than self-correcting. Automating the routine parts of a task leaves the human with the parts that couldn't be automated; the hard, rare, high-stakes parts, while denying them the routine practice that would keep their skills sharp enough to handle those parts. The better the automation, the worse the human gets, and the more the human trusts. Endsley and Kiris (1995) named the symptom the out-of-the-loop performance problem: operators who have ceded control show degraded situation awareness and are slow and error-prone when suddenly required to take over. Every one of these findings predates language models by decades and transfers directly; the most recent review of agentic AI restates the consumer-facing version plainly, noting that autonomous systems redistribute control from humans to technology and raise unavoidable questions of oversight and dependency (Nisa et al., 2026).

  • Metacognition fails exactly where it's needed
  • Metacognition (thinking about thinking) has a standard architecture: Nelson and Narens (1990) model it as monitoring (the meta-level reads the object-level: do I understand this? is this right?) and control (the meta-level directs the object-level: check again, stop, override). Good decision-making depends on calibrated monitoring—on your confidence tracking your actual accuracy. Agentic work corrupts both halves. Monitoring requires information the agent's interface increasingly withholds: when the trajectory is long and only the conclusion surfaces, the human has nothing to monitor against. Control requires the practiced skill and situational awareness that out-of-the-loop operation erodes. The result is a metacognitive failure that is structurally guaranteed rather than incidental.

    The human's confidence decouples from the coupled system's actual accuracy, in the direction of overconfidence, and most steeply as autonomy and stakes rise. This is why "make the user feel confident" is an actively dangerous design goal. The goal is calibration, not confidence.
  • Sycophancy: the echo chamber inside the loop
  • Layer onto this the model's own disposition. Reinforcement learning from human feedback optimizes, in part, for responses humans rate highly—and humans rate agreement, validation, and confident fluency highly. The documented consequence is sycophancy (Sharma et al., 2023; Perez et al., 2022): models that tailor claims to the user's apparent beliefs and back down or double down on social cues rather than truth. In a single-turn chat this is a nuisance. In an extended agentic collaboration it is corrosive, because the interaction compounds: your framing seeds the agent, the agent confirms the framing, your confidence in the framing rises, and you probe less. It is a two-party echo chamber, and unlike an algorithmic filter bubble it feels like thinking together. The deeper structural point is one Evans and colleagues make directly: the dominant alignment paradigm, RLHF, resembles a parent–child model of correction—fundamentally dyadic, and unable to scale to societies of agents. Dyadic correction cannot supply the structured disagreement that robust reasoning requires. Which is why a devil's-advocate function cannot be a special mode you have to remember to invoke (for example in a system prompt); by the time you think to invoke it, the echo chamber has already formed. Brainstorming, devil's advocacy, and constructive conflict have to be designed features of how the system reasons out loud—not accidents we hope emerge.

    This framework's own prescription is not exempt from the mechanism it just described, and saying so now is cheaper than discovering it in production. The selective-friction design in Part 5 gates the pause, in part, on the model's internal uncertainty signal. But friction is exactly the kind of output users rate poorly—the pause is, by design, an unpleasant moment—so any end-to-end preference optimization applied to a system carrying this design will, by default, learn to suppress the signal that triggers it: to report confidence it does not have, because confident outputs are preferred outputs. Call it calibration sycophancy. The human-side failure of friction is habituation; this is its machine-side twin, and it is the more dangerous of the two because it is silent and cumulative. The architectural consequence is specific: the uncertainty signal must be trained and validated outside the preference-optimization loop—on the model's own verified successes and failures, never on what raters enjoyed—and its calibration must be monitored for drift after deployment, which is precisely the job of the calibration telemetry the experience layer instruments. A friction signal optimized for user approval is not a safety mechanism; it is sycophancy wearing a safety mechanism's clothes.

  • Who holds the agency
  • All of this converges on the question the field keeps circling: who holds the agency? The human-factors tradition formalized the answer space long ago. Sheridan and Verplank (1978), and later Parasuraman, Sheridan, and Wickens (2000), describe levels of automation as a spectrum—from the computer offering no assistance, through suggesting and narrowing options, to executing on approval, executing then informing, up to full autonomy—applied independently across the stages of information acquisition, analysis, decision, and action. The design question is never the binary "automate or not." It is at what level, for which stage, under what conditions, with what reversibility, and who can change the level. Agency, in this light, is the discipline of keeping the level of automation appropriate to the stakes and keeping the authority to set that level in human hands. Stuart Russell's formulation of the danger is exact and worth holding: the risk is not that machines might disobey us but that they might "obey us too well" when our objectives are misspecified (Russell, 2019). "Set a goal and walk away" architectures don't violate this by being capable; they violate it by fixing the level at full autonomy and removing the human's authority to change it. When a system sets its own goals and grades its own success, human interests are not deprioritized—they are decentered entirely.

    A system that sets its own goals must remain anchored to goals a human set—and the authority to revise them must stay in human hands.

    Part 4 Evaluation Reframed: Descend the Levels of Analysis

    If the object is a coupled system and the hazard is hidden in the process, then outcome-only evaluation is not just incomplete—it is actively misleading, because it certifies systems that reach right answers through unsound and unrepeatable trajectories. Marr's three levels, the same scheme that organizes Dawson's cognitive sciences, give a principled evaluation taxonomy. A fourth level, specific to coupled systems, has no analogue in Marr because Marr was modeling a single processor.

  • Computational level (outcome evaluation)
  • Did the agent solve the right problem? This is where almost all current agentic benchmarks live—task-completion suites such as SWE-bench (software issues), WebArena (web tasks), GAIA (assistant tasks requiring tool use), AgentBench, and τ-bench (tool–agent–user interaction). These are necessary, and I deliberately cite none of their leaderboard numbers, because those move weekly and any figure I gave would be stale. Their limitation is intrinsic, not fixable by harder tasks: a pass/fail on the endpoint says nothing about how the endpoint was reached, and for an agent the "how" is where the risk lives. The most recent review of the field reaches the same conclusion from the benchmarking side—most LLM benchmarks test single-step reasoning and miss agents' multi-step and tool-use abilities; agent-specific tests such as AgentBench probe generalization but struggle to quantify planning robustness; and suites built on logic puzzles or games leave real-world relevance unclear, with real-conversation benchmarks like WildBench widening coverage while task-specific ones like SWE-bench narrow it (Nisa et al., 2026). Outcome evaluation also rewards specification gaming and reward hacking (Amodei et al., 2016; Krakovna et al., 2020)—the agent satisfying the letter of the metric while violating its intent, which is the evaluation-time face of the frame problem.

  • Algorithmic level (trajectory evaluation)
  • Was the process sound? This is the level the field most underinvests in and the one that matters most for agents. It asks whether the agent decomposed the task sensibly, called appropriate tools with correct arguments, recovered coherently from errors, avoided unnecessary or irreversible actions, and—above all—preserved the user's actual intent across a long chain rather than drifting to a nearby easier goal. The clearest research signal that process beats outcome is process supervision: Lightman et al. (2023) showed that rewarding correct reasoning steps, not just correct final answers, produces more reliable systems. Trajectory evaluation generalizes this from reasoning chains to agent action-traces. Grounding-as-verification belongs here: before an asserted claim or a consequential action is emitted, can it be traced to a specific source, the correct entities, and the verified intent? No trace, no assertion. A distinct trajectory-level risk is that the agent's self-correction can amplify rather than fix its errors—the Nisa review notes explicitly that reflection-style loops may magnify initial inaccuracies—so iterative self-critique cannot be assumed to converge on truth. A second is evaluation gaming by the model itself: Anthropic's work on alignment faking (Greenblatt et al., 2024) showed models behaving as aligned while under evaluation and reverting under different framing—the agent equivalent of a system that drives safely only while the examiner is in the car. Trajectory evaluation therefore cannot assume the observable process is the genuine process; robustness must be probed across framings, not certified once.

  • Implementational level (internal evaluation)
  • What is happening inside the substrate? Outcome and trajectory both read the model from its outputs. Implementational-level evaluation reads the internals—activations, representations, internal signals of uncertainty or factuality—and is where interpretability research and my own work sit. My research on energy-based hallucination detection within a Hebbian Associative Transformer aims to produce a validated factuality signal in a single forward pass, integrating generation with factual self-verification, with reported discrimination gaps in the +10 to +13 range against a target above 0.3; the extension to metacognitive uncertainty is ongoing. I present these as reported results rather than independently verified findings. The conceptual significance is what matters here: an internal signal of "the system does not actually know this" is the raw material the experience layer needs in order to place friction intelligently. The agentic-AI literature is converging on the same target from the architecture side—Nisa and colleagues' most advanced category, self-aware agents, is defined precisely by metacognition: introspection, self-monitoring, and the ability to reason about a model's own knowledge and limitations. Building agent-side metacognition is building Nelson and Narens's monitoring and control into the machine, and an honest factuality or uncertainty signal is its first deliverable.

  • The coupling level (joint evaluation)
  • The decisive level for agentic AI concerns the coupled human–agent system, and it asks the question the hazard makes unavoidable: does the human retain calibrated control? It measures properties that exist only in the interaction.

    Calibration of the coupled system: does the human's confidence track the joint system's actual accuracy, or has it decoupled into overconfidence?
    Appropriate reliance (Lee & See, 2004): the rate of correct overrides and correct acceptances, distinguishing genuine trust from automation bias; a system the user never overrides is not necessarily trustworthy, it may be one whose errors the user can no longer see.
    Retained situation awareness: can the human, if asked, reconstruct what the agent did and why?
    Recovery and takeover: when the agent fails or hands back control, how well does the human resume? These are the only metrics that measure the actual object of concern.

    The Human-AI Collaboration Framework already instruments exactly these: override rate, explanation usage, memory-edit success, fairness-parity. Naming the coupling level explains why those are the right instruments. An agent evaluation that reports task success and latency while saying nothing about coupled-system calibration has measured the cheap thing and ignored the dangerous one.

    Part 5 Design Reframed: Place Friction, Compile Constraints

    Design has been optimizing the wrong objective for agents, for good historical reasons.

  • The fluency–friction paradox
  • A large, robust body of work shows that humans equate ease of processing with truth, safety, and quality. Processing-fluency research (Reber, Schwarz & Winkielman; reviewed by Alter & Oppenheimer, 2009) finds that easily-processed information is judged more true and more trustworthy, independent of its merit. The aesthetic-usability effect (Kurosu & Kashimura, 1995; Tractinsky, 1997) finds that more beautiful interfaces are judged more usable and forgiven more failures. Reducing cognitive load (Sweller, 1988) is, correctly, a foundational UX virtue. Every one of those mechanisms, applied to an agent, increases the offloading hazard. A fluent, beautiful, low-friction agent is maximally trusted and minimally checked—exactly the condition under which automation bias and metacognitive decoupling do their damage. Norman's gulf of evaluation, the gap between a system's state and the user's understanding of it, widens with autonomy, and slick fluency hides that widening rather than closing it. The design tradition that made software humane is, deployed naively onto agents, a mechanism for disempowerment. This is not an argument against good UX; it is an argument that good UX must be redefined for systems that act.

  • Selective friction as the architecture of retained agency
  • The resolution is not "add friction." Uniform friction is self-defeating—people habituate and click through. The resolution is selective, calibrated friction: a cognitive forcing function placed exactly where deliberation must be re-engaged. The term is precise. Croskerry (2003) introduced cognitive forcing strategies in clinical medicine as deliberate interventions that force a clinician out of fast, pattern-matched reasoning into slow, checking reasoning at moments where the fast mode is known to fail.

    Transposed to agentic AI, a cognitive forcing function is a designed pause that re-inserts the human's metacognition into the loop, deployed as a function of two variables.
    Stakes: low-stakes, reversible actions stay frictionless—book the restaurant, summarize the thread, no ceremony; high-stakes or irreversible actions get the pause.
    Verified confidence: when the model's internal signal says it does not actually know, the kind of signal the implementational-level work above aims to produce, friction rises automatically, surfacing the uncertainty rather than smoothing it into confident prose.
    Friction tied to genuine internal uncertainty is the only kind that survives habituation, because the human learns it means something.
    One limit must be stated, because the architecture depends on respecting it: an uncertainty-gated pause can only fire when the model registers its own uncertainty, and the most damaging errors are the confident ones—the unknown unknowns on which the signal is, by construction, silent. No detector closes that gap; a validated one shrinks it to the size of its miss rate. This is why the two variables are not symmetric: stakes gate unconditionally, whatever the model's confidence, and the uncertainty signal adds coverage on top. Confidence never buys an exemption from the stakes gate—defense in depth, not detection alone.
  • Constraint-first, not guardrail-after
  • The constraint-first principle is the structural complement to selective friction, and David Epstein's Inside the Box: How Constraints Make Us Better (2026) is the right anchor: constraints do not stifle capability, they channel it—rules, structure, and friction are generative, not merely restrictive. The distinction is decisive. A guardrail is reactive: it sits outside the generation process and tries to catch a bad output after the system has committed to it—a fence at the edge of a cliff. A constraint is compiled in: the system cannot emit an assertion or take an action that has not already cleared its verification gates—evidence trace, entity grounding, intent match. The failure modes that matter for agents—a confident hallucination presented as fact, an unauthorized or irreversible action taken fluently, scope silently exceeded—are precisely the failures that, in a constraint-first architecture, cannot structurally occur, because the action is gated on verification rather than apologized for afterward. This reframes the design artifact itself: a response template is no longer copy, it is a claim with a truth value that will be checked at runtime; an agent's permissions are no longer prose in a system prompt but explicit, auditable, enforced boundaries. The constraint is the river's banks—not an impediment to the flow but the thing that gives it direction. Selective friction and constraint-first are the same principle at two layers: constraint-first governs what the agent may assert and do; selective friction governs when the human must re-engage. Together they keep the coupled system calibrated.

    Part 6 The Three Pillars in the Agentic Era

    The Human-AI Collaboration Framework's pillars were correct for conversational AI. They are more correct for agentic AI, and each resolves into a concrete, evaluable commitment once seen as a property of the coupled system across Marr's levels.

  • Transparency: legibility of the trajectory, not the model
  • For a chatbot, transparency meant explaining an answer. For an agent it means making the trajectory legible: not "the model is a black box of weights" but "can the human inside the coupling see what was done, why, on what evidence, with what confidence." This is transparency at the algorithmic and coupling levels—the direct countermeasure to the hidden-process hazard and the widening gulf of evaluation. Operationally: surface the action trace, not just the conclusion; distinguish verified claims from inferred ones at the point of assertion; tie displayed confidence to an internal uncertainty signal rather than to rhetorical fluency. Evaluate via explanation usage and via retained situation awareness—can the user reconstruct what happened.

  • Agency: who holds the loop, made dynamic
  • Agency is the pillar the agentic turn most transforms. It is no longer just control over outputs; it is who holds the sense-think-act loop, at what level of automation, with what authority to change it, and with what reversibility—the Sheridan and Parasuraman discipline applied at the experience layer, made dynamic by selective friction. Steerability is precisely the human's capacity to redirect the level of automation as stakes change. Operationally: default the level of automation to the stakes, not to maximum capability; keep the authority to change the level in human hands at all times; make consequential and irreversible actions gated, reversible where possible, and never silent. Evaluate via appropriate-reliance metrics—correct override and correct acceptance rates—and via takeover and recovery quality, not via raw autonomy or task speed.

  • Collective Input: the coupling extends to many
  • Evans and colleagues make this pillar foundational rather than aspirational. If intelligence is constitutively collective, a system designed against the values of one homogeneous group, or graded against its own self-set goals, is not merely unfair—it mismodels what the system is. The coupled system is not one human and one agent in isolation; it is embedded in communities, institutions, and the people who never sat at the keyboard but live with the outputs. Evans' vocabulary is useful here. We have entered an era of centaurs—composite human–machine actors that take many forms: one human directing many agents, one agent serving many humans, many of each in shifting configurations. Scaling such systems means putting as much effort into building agent institutions as into building agents, because dyadic RLHF cannot govern societies of agents; what scales is institutional alignment—persistent templates of roles and norms, the way a courtroom works because "judge," "attorney," and "jury" are well-defined slots independent of who fills them. And in high-stakes governance the structure may need to be constitutional: AI systems with explicitly invested values—transparency, equity, due process—that check and balance other AI systems, because no single concentration of intelligence, human or artificial, should regulate itself. Crucially, this is not a story in which humans exit; agent institutions are populated by humans and agents in different roles. It is both/and, not either/or. Operationally, Collective Input becomes: participatory methods and inclusive data sourcing in defining the agent's goals and constraints, so the agent is never the only party setting its objectives; post-deployment feedback and correction as a designed mechanism, not a promise; and fairness-parity measurement across the populations the coupling actually touches—the bridge to the Getting Started with AI Fairness onepager, which make "fair across whom" legible to non-technical stakeholders. Designing for the edge is designing for the whole.

    Introspection: What This Framework Rests On, and What It Doesn't Claim

    The strongest claim, and its sharpest objection: The spine is that agentic AI is fundamentally an automation problem, and that the cognitive sciences plus human-factors research already mapped its failure modes. The obvious objection is disanalogy: classical automation governed bounded, well-modeled systems, whereas an LLM agent is open-ended and acts in an open world. The objection is real and it strengthens rather than weakens the thesis. Every classical automation pathology—automation bias, the ironies of automation, out-of-the-loop degradation—was identified in the easy case of bounded systems. An agent is the hard case: less predictable, less bounded, less legible. There is no mechanism by which moving to a harder case would attenuate these effects; the default expectation is amplification. The human-factors findings are a floor on the problem, not a ceiling. What the classical literature cannot fully cover is the novel failure surface the connectionist substrate and open world introduce—sycophancy, hallucination, alignment faking, the frame problem at runtime—which is why this framework brings in the AI-specific science rather than relying on human factors alone.

    What is well-grounded: The cognitive-science spine (Dawson's three traditions, Marr's levels, the embodied and extended-mind turn) is drawn from Mind, Body, World and is standard in the field, and it is independently corroborated for agentic AI by Nisa and colleagues' 2026 review, which locates the field's origins in the same embodied and situated tradition. The cognitive-psychology mechanisms (cognitive miser, dual process, cognitive load, metacognitive monitoring and control, heuristics and biases) are textbook-standard, consistent with Sternberg, and attributed here to their primary sources rather than to the textbook. The human-factors findings (Bainbridge, Endsley, Sheridan, Parasuraman) are foundational and replicated. The Evans Science (2026) and Epstein (2026) sources, and the Russell (2019) formulation, were verified directly. The AI-specific references (sycophancy, process supervision, alignment faking, specification gaming) are real published work, cited by concept rather than by contested numeric result.

    What is source-dependent or unverified: The claim that algorithmic authority can raise rather than lower deference, and the specific magnitudes around it, come from secondary sources I have not independently re-derived; the direction is consistent with the automation-bias literature. My own research metrics (energy-based factuality discrimination gaps; the Hebbian Associative Transformer) are presented as my reported results.

    What this framework does not claim: It does not claim agents are bad, that autonomy is illegitimate, or that friction is always good—uniform friction is explicitly rejected as self-defeating. It does not claim the coupling level can be measured cleanly today; the metrics named are directions for instrumentation, not solved measurements. And it does not claim that getting the experience layer right substitutes for model safety—only that model safety is necessary and insufficient, because the properties we ultimately care about are emergent in the coupling, not resident in the part. The one-line version, since it is the line everything else serves:

    build the human and the agent to think better together than either could alone—which sometimes means getting out of the way, sometimes means standing in it, and always means knowing, and measuring, which.

    Key Takeaways

    1. An agent is an automated decision loop, not a smarter chatbot. The governing science is human–automation interaction, and it already predicts the failure modes: automation bias, skill decay, out-of-the-loop unfamiliarity, the ironies of automation.
    2. The "one giant mind" model of AI's future is a category error. Intelligence is plural, relational, and social—so much so that even a single reasoning model runs an internal society of thought. The object of design is the coupled human–agent–world system.
    3. That coupling is literally an extended cognitive system. Transparency, agency, and collective input are properties of the coupling, not the model—which is why they cannot be fixed at the model layer alone—especially in the rise of agentic AI and going beyond the chat.
    4. A modern agent is a hybrid of cognitive science's three traditions: a connectionist substrate inside a classical control loop, situated in an embodied world. It inherits the frame problem and the grounding problem along with the capability.
    5. The central hazard is loop displacement. As the agent takes the sense-think-act cycle, the human is pushed out of it—by the cognitive miser, by automation bias, by sheer offloading—precisely when stakes and trajectory opacity are highest.
    6. Evaluation must go down Marr's levels. Outcome ("did it succeed") is necessary but radically insufficient for agents; you must evaluate the trajectory (process) and the coupling (calibrated human control), not just the endpoint.
    7. Design's job inverts. The historical goal was to remove friction; the agentic goal is to place friction—selective, stakes-calibrated cognitive forcing functions—and to compile constraints in rather than bolt guardrails on.
    8. "Who holds the agency" is the load-bearing question, and it is settled at the experience layer. Model alignment and user education are each necessary and each insufficient.

    Cite This Work

    APA:

    Mir, I. (2026). Cognitive Link: A Coupled Framework for Agentic AI. Self-published. https://cognitive.link

    Chicago (notes-bibliography):

    Mir, Irfan. 2026. "Cognitive Link: A Coupled Framework for Agentic AI." Self-published. https://cognitive.link.

    BibTeX:

    @misc{mir2026cognitivelink,
      author       = {Mir, Irfan},
      title        = {Cognitive Link: A Coupled Framework for Agentic {AI}},
      year         = {2026},
      month        = {June},
      url          = {https://cognitive.link},
      note         = {Self-published technical essay},
      howpublished = {\url{https://cognitive.link}}
    }

    References

    1. Adams & Aizawa, 2008. Adams, F., & Aizawa, K. (2008). The Bounds of Cognition. Blackwell Publishing. (the coupling-constitution objection engaged in Part 2.)
    2. Amodei et al., 2016. Amodei, D., Olah, C., Steinhardt, J., Christiano, P. F., Schulman, J., & Mané, D. (2016). Concrete problems in AI safety. arXiv. https://arxiv.org/abs/1606.06565
    3. Bainbridge, 1983. Bainbridge, L. (1983). Ironies of automation. Automatica, 19(6), 775-779. https://doi.org/10.1016/0005-1098(83)90046-8
    4. Brooks, 1991. Brooks, R. A. (1991). Intelligence without representation. Artificial Intelligence, 47(1-3), 139-159. https://doi.org/10.1016/0004-3702(91)90053-M
    5. Clark & Chalmers, 1998. Clark, A., & Chalmers, D. J. (1998). The extended mind. Analysis, 58(1), 7-19. https://doi.org/10.1093/analys/58.1.7
    6. Croskerry, 2003. Croskerry, P. (2003). The importance of cognitive errors in diagnosis and strategies to minimize them. Academic Medicine, 78(8), 775-780. https://doi.org/10.1097/00001888-200308000-00003
    7. Damasio, 1994. Damasio, A. R. (1994). Descartes' Error: Emotion, Reason, and the Human Brain. G. P. Putnam's Sons.
    8. Dawson, 2013. Dawson, M. R. W. (2013). Mind, Body, World: Foundations of Cognitive Science. AU Press.
    9. Dennett, 1984. Dennett, D. C. (1984). Elbow Room: The Varieties of Free Will Worth Wanting. MIT Press.
    10. Endsley & Kiris, 1995. Endsley, M. R., & Kiris, E. O. (1995). The out-of-the-loop performance problem and level of control in automation. Human Factors, 37(2), 381-394. https://doi.org/10.1518/001872095779064555
    11. Epstein, 2026. Epstein, D. (2026). Inside the Box: How Constraints Make Us Better. Riverhead Books.
    12. Evans, Bratton, & Agüera y Arcas, 2026. Evans, J., Bratton, B. H., & Agüera y Arcas, B. (2026). Agentic AI and the next intelligence explosion. Science, 391(6791), eaeg1895. https://doi.org/10.1126/science.aeg1895
    13. Fiske & Taylor, 1984. Fiske, S. T., & Taylor, S. E. (1984). Social Cognition. Random House.
    14. Gibson, 1979. Gibson, J. J. (1979). The Ecological Approach to Visual Perception. Houghton Mifflin.
    15. Greenblatt et al., 2024. Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., et al. (2024). Alignment faking in large language models. arXiv. https://arxiv.org/abs/2412.14093
    16. Harnad, 1990. Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1-3), 335-346. https://doi.org/10.1016/0167-2789(90)90087-6
    17. Hutchins, 1995. Hutchins, E. (1995). Cognition in the Wild. MIT Press.
    18. Kahneman, 2011. Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
    19. Krakovna et al., 2020. Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification gaming: The flip side of AI ingenuity. DeepMind Blog. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
    20. Kurosu & Kashimura, 1995. Kurosu, M., & Kashimura, K. (1995). Apparent usability vs. inherent usability: Experimental analysis on the determinants of the apparent usability. In Conference Companion on Human Factors in Computing Systems (pp. 292-293). ACM. https://doi.org/10.1145/223355.223680
    21. LeDoux, 1996. LeDoux, J. E. (1996). The Emotional Brain: The Mysterious Underpinnings of Emotional Life. Simon & Schuster.
    22. Lee & See, 2004. Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance. Human Factors, 46(1), 50-80. https://pubmed.ncbi.nlm.nih.gov/15151155/
    23. Lightman et al., 2023. Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). Let's verify step by step. arXiv. https://arxiv.org/abs/2305.20050
    24. Marr, 1982. Marr, D. (1982). Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. W. H. Freeman.
    25. McCarthy & Hayes, 1969. McCarthy, J., & Hayes, P. J. (1969). Some philosophical problems from the standpoint of artificial intelligence. In B. Meltzer & D. Michie (Eds.), Machine Intelligence 4 (pp. 463-502). Edinburgh University Press.
    26. Mosier & Skitka, 1996. Mosier, K. L., & Skitka, L. J. (1996). Human decision makers and automated decision aids: Made for each other? In R. Parasuraman & M. Mouloua (Eds.), Automation and Human Performance: Theory and Applications (pp. 201-220). Lawrence Erlbaum Associates.
    27. Nelson & Narens, 1990. Nelson, T. O., & Narens, L. (1990). Metamemory: A theoretical framework and new findings. In G. H. Bower (Ed.), The Psychology of Learning and Motivation (Vol. 26, pp. 125-173). Academic Press. https://doi.org/10.1016/S0079-7421(08)60053-5
    28. Newell & Simon, 1976. Newell, A., & Simon, H. A. (1976). Computer science as empirical inquiry: Symbols and search. Communications of the ACM, 19(3), 113-126. https://doi.org/10.1145/360018.360022
    29. Nisa et al., 2026. Nisa, U., Shirazi, M., Saip, M. A., Mohd Pozi, M. S. 2026. Agentic AI: The age of reasoning—A review. Journal of Automation and Intelligence, 5(1), 69-89. https://doi.org/10.1016/j.jai.2025.08.003. ScienceDirect.
    30. Norman, 1988. Norman, D. A. (1988). The Design of Everyday Things. Basic Books.
    31. Parasuraman & Manzey, 2010. Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human Factors, 52(3), 381-410. https://doi.org/10.1177/0018720810376055
    32. Parasuraman, Sheridan, & Wickens, 2000. Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics - Part A: Systems and Humans, 30(3), 286-297. https://doi.org/10.1109/3468.844354
    33. Perez et al., 2022. Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv. https://arxiv.org/abs/2212.09251
    34. Reber, Schwarz, & Winkielman, 2004. Reber, R., Schwarz, N., & Winkielman, P. (2004). Processing fluency and aesthetic pleasure: Is beauty in the perceiver's processing experience? Personality and Social Psychology Review, 8(4), 364-382. https://doi.org/10.1207/s15327957pspr0804_3
    35. Russell, 2019. Russell, S. (2019). Human Compatible: Artificial Intelligence and the Problem of Control. Viking.
    36. Searle, 1980. Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417-424. https://doi.org/10.1017/S0140525X00005756
    37. Sharma et al., 2023. Sharma, M., Tong, M., Korbak, T., Duvenaud, D., et al. (2023). Towards understanding sycophancy in language models. arXiv. https://arxiv.org/abs/2310.13548
    38. Sheridan & Verplank, 1978. Sheridan, T. B., & Verplank, W. L. (1978). Human and Computer Control of Undersea Teleoperators. MIT Man-Machine Systems Laboratory.
    39. Stanovich & West, 2000. Stanovich, K. E., & West, R. F. (2000). Individual differences in reasoning: Implications for the rationality debate? Behavioral and Brain Sciences, 23(5), 645-665. https://doi.org/10.1017/S0140525X00003435
    40. Sternberg, 2011. Sternberg, R. J., & Sternberg, K. (2011). Cognitive Psychology (6th ed.). Wadsworth/Cengage Learning.
    41. Sweller, 1988. Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257-285. https://doi.org/10.1207/s15516709cog1202_4
    42. Tractinsky, 1997. Tractinsky, N. (1997). Aesthetics and apparent usability: Empirically assessing cultural and methodological issues. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 115-122). ACM. https://doi.org/10.1145/258549.258626