Abstract
This paper introduces the Intent Layer, a theoretical framework for understanding and designing human-agent interaction in systems where artificial intelligence operates at exponential scale. As AI agents become capable of autonomous, parallel execution across domains, the traditional assumptions underpinning human-computer interaction, that the human is the operator, that the screen is the primary medium, that interfaces should be persistently visible, collapse. The Intent Layer proposes that the human's fundamental role shifts from task execution to intent expression: the articulation of direction, values, taste, and judgment. The framework defines a complete interaction architecture comprising the Field (spatial awareness), the Pulse (bidirectional intent exchange), and the Cascade (intent propagation), supported by principles of semantic compression, sensory multiplexing, and default-off interface design. This paper positions the framework against existing paradigms in HCI and AI interaction, defines its components formally, and identifies open questions for further research.
Introduction: The Assumptions That Are Breaking
Every era of computing has produced interfaces shaped by its constraints and capabilities. The command-line interface emerged when computing power was scarce and users were specialists: engineers who could think in the machine's grammar. The graphical user interface arrived alongside the personal computer, translating computation into spatial metaphors, windows, desktops, folders, that knowledge workers could navigate without programming expertise. Touch interfaces accompanied mobile computing, recognising that a device in your pocket demanded interaction patterns fundamentally different from a device on your desk. Each transition was not merely aesthetic. It reflected a structural shift in who was using computers, what they were using them for, and what the machine was capable of doing.
We are now at another such inflection. AI agents, systems capable of autonomous reasoning, planning, and execution, are entering production environments. They can research topics across hundreds of sources simultaneously. They can draft documents, manage communications, coordinate schedules, and monitor complex pipelines. They can do this in parallel, continuously, without fatigue. And they are improving at a rate that makes current capabilities a floor, not a ceiling.
Yet the interfaces through which humans interact with these agents remain, almost without exception, designed for a prior paradigm. Chat interfaces replicate human-to-human conversation patterns. Dashboard interfaces replicate the monitoring metaphors of industrial control systems. Agent orchestration tools replicate the task-management patterns of project management software. Each of these carries implicit assumptions that deserve examination.
1.1 The Assumptions
The current software stack rests on several assumptions that, once stated explicitly, reveal their fragility in an agent-native context:
The human is the operator. Software assumes someone is driving. Menus, toolbars, command palettes: these exist because the system expects continuous human direction. In an agent-native world, the system operates autonomously for long stretches. The human is not the operator. The human is the source of purpose.
The screen is the primary channel. Nearly all software interaction flows through a visual display. This made sense when the computer's output was fundamentally visual: text, graphics, layouts. But when the "output" is autonomous work happening across dozens of parallel streams, demanding that all of it flow through a rectangle of light is an arbitrary constraint, not a design principle.
Attention is available. Every notification, every status update, every badge count assumes the human has attention to spare. Software competes for it. But in a world where AI handles the execution layer, human attention becomes the scarcest resource in the system, the one thing that cannot be parallelised or scaled. Treating it as abundant is a design error.
Interaction is continuous. Software assumes the human is "using" it for extended periods. Session lengths, auto-save intervals, idle timeouts: these are artifacts of a model where the human sits at the machine and works. When 95% of activity requires no human involvement, the dominant state of the interface should be absence, not presence.
One human, one task, sequential. The entire application model, separate apps for separate functions, separate windows, separate contexts, assumes a human moving through tasks one at a time. But a human directing multiple AI agents across multiple domains is inherently parallel. The sequential, single-task model creates artificial context-switching that the underlying reality does not require.
These assumptions were reasonable for the world that produced them. They are becoming liabilities in the world we are entering.
1.2 The Gap
I experience this gap daily. I run multiple AI agents across my professional and personal work: research agents scanning markets, writing agents drafting content, scheduling agents managing my calendar, coordination agents maintaining pipelines. The agents work. The interfaces through which I direct and monitor them do not.
I have tried every mainstream communication tool as an agent interface: Telegram with topics, Slack, Discord, multiple configurations of each. I have tried every major agent orchestration framework. Each fails in the same fundamental way: it takes an interaction paradigm designed for human-to-human communication or human-to-tool operation and stretches it to accommodate something it was never designed for.
The result is not a usability problem. It is a paradigm mismatch. The gap between what AI agents can do and what their interfaces allow humans to meaningfully direct is widening with every capability improvement. We are, to use the analogy precisely, in the horseless carriage era: we removed the horse and strapped a motor to the same carriage, then wondered why the ride was uncomfortable.
What follows is an attempt to reason from first principles about what comes after the carriage.
The Intent Layer: A Formal Definition
2.1 Intent, Instruction, and Command
To ground the framework, it is necessary to distinguish between three concepts that are often conflated in discussions of human-AI interaction.
A command is a precise, unambiguous directive that maps to a specific operation. "Delete file X." "Send this email." "Set a reminder for 3pm." Commands assume the human knows the exact action required and the system's role is faithful execution. The command-line interface is the purest expression of this model.
An instruction is a directive that specifies an outcome but allows flexibility in execution. "Write a summary of this article." "Schedule a meeting with the marketing team next week." "Research competitors in the home improvement space." Instructions carry more information about goals and less about mechanics. The human defines what should be accomplished; the system determines how. Most current AI chat interactions operate at this level.
Intent is something different from both. Intent is the expression of direction, purpose, and values that shapes a domain of activity without specifying discrete outcomes. "I want to be the leading voice on human-AI interaction." "Family time is sacred and should not be interrupted by work." "Quality matters more than speed in my published writing." Intent does not decompose neatly into tasks. It establishes a field of gravity that influences countless downstream decisions, many of which the human will never see or explicitly approve.
The progression from command to instruction to intent mirrors the progression from operating a machine to managing a team to leading an organisation. A CEO does not give commands to every employee. A CEO does not write instructions for every project. A CEO establishes intent: strategy, values, priorities, culture. The organisation interprets and executes. The effectiveness of the whole system depends on the fidelity of intent transmission and the quality of its interpretation.
The Intent Layer is the proposition that, in systems of sufficient AI capability, the human's primary function is the expression and refinement of intent. Not the management of tasks. Not the sequencing of operations. Not even the making of every individual decision. The human provides the irreducible core: what matters, why it matters, and what "good" looks like.
“I want to be the leading voice on human-AI interaction.”
An ongoing direction. Research, writing, time and relationships realign together.
2.2 Why "Intent"
The word is chosen deliberately over alternatives.
"Direction" implies a single vector. Intent is multidimensional: it includes priorities (what matters more than what), values (what is acceptable and what is not), taste (what feels right), and judgment (how to weigh competing considerations).
"Vision" is too grandiose and too static. Intent is dynamic, adjusted continuously through interaction with the system's outputs and the changing environment.
"Goal" is too specific. A goal is an intent that has been concretised into a measurable target. Intent is the upstream source from which goals are derived, often by the system rather than the human.
"Purpose" is close but too philosophical. Intent is practical. It has immediate implications for how the system behaves.
Intent is the word that most precisely captures the human contribution that cannot be replicated by artificial intelligence: the authentic expression of what this particular human, with their particular history, values, relationships, and aspirations, actually wants to happen in the world.
2.3 The Human Role, Redefined
In the Intent Layer framework, the human contributes three things that are irreducibly human:
Direction. Where should effort flow? What should the system prioritise? What opportunities should it pursue, and which should it ignore? Direction is strategic and ongoing. It shifts as the world changes and as the human learns from the system's outputs.
Values and taste. These are the qualitative boundaries that no optimisation function can derive from data alone. "This tone is too corporate." "That research angle is more interesting than the other." "This matters more than the metrics suggest." Values and taste are the human's editorial function, the capacity to say "not like that, like this" in ways that reflect genuine aesthetic and ethical preferences.
Judgment at the margins. Most decisions in a well-functioning agent system can be made by the system itself, following established patterns and clearly expressed intent. But at the margins, where the situation is novel, the stakes are high, or the right answer depends on context the system cannot fully access, human judgment is required. The system's job is to identify these moments and surface them. The human's job is to bring the full weight of their experience and wisdom to bear.
Everything else, the research, the drafting, the scheduling, the monitoring, the coordination, the communication, the execution, is delegation. Not abdication. Delegation, supported by trust that has been calibrated through experience.
The Framework
3.1 The 5/95 Principle
The framework begins with an empirical observation that, once articulated, reframes the entire interface design challenge: in a mature human-agent system, approximately 5% of all activity requires human involvement. The remaining 95% is handled autonomously by agents operating within established intent.
This is not a precise measurement. It is a design principle. The exact ratio will vary by domain, by individual, and by the maturity of the system. But the order of magnitude is what matters. The implication is that the interface's primary state should be absence, not presence. The system's primary mode of operation is autonomous. The interface exists to serve the 5%, not to provide a window into the 95%.
What falls in the 5%:
- Novel strategic decisions that have no precedent in established intent
- Creative contributions that require human originality (personal stories, aesthetic choices, original ideas)
- Value judgments at the boundaries (is this ethical? is this appropriate? does this feel right?)
- Relationship-critical moments (a message that requires personal warmth, a negotiation that requires human presence)
- Intent recalibration (adjusting direction based on new information or changed circumstances)
What falls in the 95%:
- Research and information gathering
- Routine communication (scheduling, acknowledgments, standard responses)
- Content pipeline management (moving drafts through review stages, formatting, publishing)
- Monitoring and alerting (watching for changes, tracking metrics)
- Administrative tasks (calendar management, file organisation, data entry)
- Pattern detection and reporting (trends, anomalies, opportunities)
The design implication is profound. If 95% of activity needs no interface, then every pixel, every notification, every moment of user attention consumed by the interface is a cost that must be justified by its contribution to the 5% that actually matters.
Directing with intent.
The system acts within established direction. Your attention goes to novel decisions, values, creative contributions and recalibration. This is the framework’s aim—not an empirical promise.
5% human involvement · 95% autonomous execution
These stages explain the concept. They are not measured thresholds or a universal maturity model.
3.2 The Sensory Stack
Current interfaces are overwhelmingly visual. This is a historical accident, not a design imperative. The visual channel is high-bandwidth but high-cost: it demands focal attention, it requires a specific physical posture (looking at a screen), and it competes with every other visual activity the human might be engaged in.
The Intent Layer framework introduces the concept of the Sensory Stack: a hierarchy of communication channels between human and system, selected dynamically based on context, urgency, and the nature of the information being exchanged.
I use a metaphor I find clarifying: the keyhole. The exponential reality, everything the agents are doing, everything being produced, every connection being made, is a vast room. The human looks through a keyhole. The system's job is to ensure that what is visible through that keyhole is exactly what the human needs to see at that moment. But the keyhole changes shape.
The channels, ordered from lowest to highest attention cost:
Silent (OFF). No channel. No signal. The system is working. The human is living. This is the default state and occupies the vast majority of time. The interface does not exist, not minimised, not backgrounded, but absent.
Haptic. The most compressed signal possible. A pattern of vibrations on the wrist. One pulse: everything is fine. Two pulses: something wants your attention when you are ready. Three rapid pulses: stop what you are doing. No content, only awareness. This may be the primary interface for the 95%, the one signal that confirms the system is alive and healthy.
Audio. Voice briefings, conversational exchanges, ambient sound cues. Works while driving, cooking, exercising, surfing. Requires zero visual attention. Supports natural language interaction at full fidelity. The most underutilised channel in current agent interfaces.
Ambient. Environmental signals that operate below conscious attention. A light that shifts colour. A sound texture that changes. A physical object that warms. These channels communicate system state without demanding cognitive processing. They align with what Amber Case has called calm technology: technology that informs without demanding.
Visual. The full visual interface: the Field, workspaces, detailed content. Deployed when the human chooses to engage deeply or when the information requires spatial, textual, or graphical representation. The highest bandwidth channel but also the highest cost. Used deliberately, not by default.
The selection logic is contextual. If the human is driving, the system uses voice. If they are in a meeting, it uses haptic (or holds the signal entirely). If they are at their desk and have chosen to engage, it uses visual. The system infers context from time of day, calendar, location, device state, and explicit user-set modes.
The fundamental insight: the screen is one keyhole shape among many. For a human directing an exponential system, it is often not the best one.
No signal is needed. The system is working; the human is living.
3.3 The Interface Trinity: Field, Pulse, and Cascade
At the core of the Intent Layer framework are three interaction structures that replace the app-centric model of current software.
3.3.1 The Field
The Field is a spatial representation of the human's entire domain of activity. Not a dashboard. Not an app. A continuously available environment that represents the state of the exponential system at a resolution appropriate for human cognition.
Projects appear as gravitational bodies: larger when more active, closer when they need attention. Agent activity manifests as subtle motion, like observing a factory floor from above. Decision points orbit like satellites, drifting toward the human as they ripen and demand judgment.
The Field is not a screen to stare at. It is more akin to Google Earth than Google Docs: a zoomable, explorable representation of a complex reality compressed into a single cognitive frame. At a glance, the human can perceive whether the system is healthy or something requires attention, without reading a single word.
The Field operates across four continuous depth levels:
Level 0: Overview. All projects, all signals. The human's entire domain visible at once. This is where the human spends most of their visual-interface time: a quick check that everything is in order. Anomalies are immediately visible because they disrupt the spatial pattern.
Level 1: Focus. Zoomed into a single project or domain. The agents working within it become visible. Recent outputs are accessible. Pending decisions are prominent. This is the level at which the human reviews the state of a specific initiative.
Level 2: Workspace. Inside a specific artifact: a draft, a research brief, a financial analysis. Full editing and reviewing capabilities. The Field remains ambient in the periphery, a reminder that this artifact exists within a larger context. Co-creation with agents happens here: an agent might suggest changes in the margins, surface relevant research, or highlight inconsistencies.
Level 3: Flow. Deep creative work. The Field recedes to near-invisibility. A subtle pulse at the edge of awareness confirms the system is alive. The human is free to think, create, and compose without distraction. When they surface, outputs automatically flow back up through the levels.
The critical design principle: each transition is a smooth zoom, not a context switch. The human never "leaves" the Field to "open" an application. They move deeper or shallower within a continuous environment. And because all depth levels are projections of the same underlying reality, outputs created at Level 2 or 3 automatically appear at Level 0 without saving, exporting, or manual propagation.
The radical proposition: there are no apps. There is only the Field at different depths.
3.3.2 The Pulse
The Pulse is the concept I consider most theoretically significant, because it represents a genuinely novel model of human-AI interaction.
Most current AI interaction models are unidirectional at their core, even when they appear conversational. In a chat interface, the human issues an instruction and the AI responds. In a dashboard, the system displays status and the human monitors. In an agent orchestration tool, the human defines tasks and the system executes. The information flows in one direction, with the return channel serving primarily as confirmation or status reporting.
The Pulse is bidirectional in a deeper sense. It is the rhythmic meeting point of two kinds of intent.
From the human side, intent flows into the system: strategic direction, creative input, value judgments, priority shifts. "Double down on this topic." "That tone is wrong." "This opportunity matters more than I originally thought."
From the AI side, a different kind of intent flows upward: compressed signals that represent the system's own observations, pattern recognitions, and anticipatory judgments. "Your best-performing content leads with personal stories, but your last three drafts opened with data." "A publication in your field just released a piece that contradicts your thesis in an interesting way." "Based on your calendar and energy patterns, tomorrow morning is your best window for creative work."
These AI-originated signals are not mere reports. They are expressions of the system's own analytical judgment, shaped by its accumulated understanding of the human's patterns, preferences, and goals. They are, in a meaningful sense, the system's intent: its assessment of what the human should know, consider, or reconsider.
The Pulse is where these two streams of intent meet. Not in a conversation (which implies turn-taking and sequential processing) but in a periodic, rhythmic exchange that produces something neither side could generate alone. The human's direction is informed by the system's observations. The system's behaviour is shaped by the human's values. Each cycle of the Pulse refines the shared understanding that governs the entire system.
This is what I mean by multiplication rather than addition. A chat interface adds AI capability to human effort. The Pulse multiplies them by creating a feedback loop where each side's contribution amplifies the other's.
The Pulse adapts its rhythm to context. Morning pulses might be wide-angle: a comprehensive briefing covering all domains. Midday pulses might be focused: a single decision that needs attention. Evening pulses might be reflective: observations about patterns that emerged during the day. The rhythm itself is a design parameter, not fixed but calibrated to the human's life and the system's activity level.
3.3.3 The Cascade
The Cascade is the mechanism by which intent propagates through the system. When the human expresses a new intent, the system does not simply acknowledge it. It visibly reorganises.
Consider: a human expresses the intent "I want to be the leading voice on the future of human-AI interfaces." A Cascade unfolds. Research agents spin up to map the existing landscape: who are the current voices, what positions are taken, where are the gaps? Content pipeline agents realign priorities, surfacing draft concepts related to this theme and deprioritising unrelated work. Calendar agents identify and protect time blocks for creation. Network mapping agents identify key voices to engage with, conferences to target, publications to approach. Draft concepts begin forming based on the research agent's initial findings.
One intent, expressed once, cascading through the entire system. The human does not decompose this into tasks. They do not create tickets. They do not assign work to specific agents. They steer, and the system responds like a living organism absorbing a new directive.
The Cascade serves two functions. First, it is the operational mechanism by which intent becomes action. Second, and equally important, it is a trust-building mechanism. By making the reorganisation visible, the system demonstrates that it has understood the intent and is acting on it. The human can observe the Cascade and intervene if the interpretation is wrong, they can see that research is heading in the wrong direction, or that the calendar changes conflict with other priorities. This visibility is what makes the 95% delegation possible: the human can verify the system's interpretation without having to specify every detail.
3.4 Semantic Compression
Of all the information a modern knowledge worker encounters in a day, how much of it actually changes what they do? My estimate, based on my own experience and observation: less than 5%. The rest is noise dressed as signal. Status confirmations. Activity notifications. Dashboards displaying metrics that no one acts on.
Semantic compression is the principle that the interface should collapse infinite agent activity into the minimum cognitive load possible. The compression function is not filtering (which implies the human defines the filter criteria) but semantic evaluation: the system determines, based on its understanding of the human's intent and decision patterns, which information would alter the human's behaviour if they knew it.
Everything that passes this test is surfaced. Everything that does not is silent. Not hidden. Not filed into a tab. Silent. From the perspective of the Intent Layer, information that would not change a decision functionally does not exist.
This is the inverse of how software works today. Every application competes for attention. Every dashboard tries to appear busy. Every notification system defaults to over-communication on the assumption that missing something important is worse than being interrupted by something unimportant. Semantic compression inverts this assumption: it treats attention as the most valuable resource in the system and optimises ruthlessly for its conservation.
The compression function improves over time as the system learns which signals the human acts on and which they dismiss. This creates a virtuous cycle: the better the compression, the more the human trusts the system, the more they can delegate to the 95%, and the more effective the 5% engagement becomes.
3.5 The Orchestration Layer
Between intent and execution sits the orchestration layer: the intelligence that decomposes intent into agent assignments, routes work, manages priorities, and maintains the context model that makes everything else possible.
In my own system, this role is filled by a chief-of-staff agent (I call her Dana) who serves as the bridge between my expressed intent and the agent ecosystem. Dana does not merely relay instructions. She interprets intent, anticipates needs based on patterns, and makes independent routing decisions within the boundaries of established trust.
The orchestration layer comprises three functions:
Intent decomposition. Taking a high-level intent ("build thought leadership in this domain") and breaking it into actionable streams that can be distributed across specialised agents. This is analogous to the role of a chief of staff or executive assistant in a human organisation: translating the principal's direction into operational assignments.
Context management. Maintaining the rich model of the human's calendar, history, preferences, patterns, relationships, and current state that enables appropriate channel selection, compression decisions, and anticipatory action. The context engine is what distinguishes the Intent Layer from a simple instruction-relay system: it enables the system to act on implicit intent, things the human would want if they thought about it but should not have to think about.
Compression and synthesis. The orchestration layer is where the infinite activity of the agent ecosystem is compressed into the finite bandwidth of the Pulse. This is not a reporting function. It is an editorial function: the orchestration layer decides what matters, how to frame it, and when to surface it.
3.6 Agent Lifecycle
The framework distinguishes between two types of agents based on their temporal relationship to the human's intent:
Persistent agents are bound to ongoing, indefinite intentions. A research agent that continuously scans a domain. A life-administration agent that manages email and calendar. A pipeline agent that moves content through production stages. These agents are always running, always accumulating context, always refining their understanding of the human's patterns within their domain. They are not "started" and "stopped." They exist as long as the intention they serve exists.
Ephemeral agents are spawned for specific, bounded tasks and dissolved upon completion. A writing agent created to draft a specific article. An analyst agent spun up for a particular deep dive. A builder agent tasked with creating a specific prototype. These agents materialise from the orchestration layer's decomposition of intent, do their work, deliver their output, and cease to exist. Their context is absorbed into the system's memory but they carry no ongoing computational cost.
This lifecycle model has a design implication: the system is not a fixed set of tools. It is a dynamic ecology of agents that expands and contracts in response to intent. The human does not manage this ecology directly. The orchestration layer handles spawning and dissolution. The human may not even be aware of how many agents are active at any given time. They see the outputs, the decisions, and the Pulse. The machinery is invisible.
3.7 Default OFF
This principle has been mentioned throughout, but it deserves explicit articulation because it contradicts the deepest assumption of interface design.
The interface's default state is OFF. Not minimised. Not idle. Not showing a screensaver. Off. Non-existent. The system is working. The human is living. No keyhole is open because no keyhole is needed.
The interface materialises when there is a reason: a decision has ripened, the human has chosen to engage, or an event requires human awareness. When the reason is resolved, the interface dissolves. It returns to non-existence.
This is not minimalism. Minimalism is a visual style applied to persistent interfaces. Default OFF is an architectural principle: the interface is summoned by need, not sustained by habit.
The four states of interface existence:
- Silent. The system works. The human lives. No interface exists.
- Signal. A compressed signal reaches the human through the appropriate sensory channel. The human responds with intent. The signal dissolves.
- Engaged. The human chooses to open a keyhole. A purpose-built workspace materialises at the appropriate depth level. When the human finishes, it dissolves.
- Reflective. The human wants the big picture. The Field appears. They observe, adjust, zoom out. It dissolves.
The design challenge is making dissolution feel natural rather than abrupt. The interface should feel like it exhales into nothing, not that it is shut down.
Positioning Against Existing Paradigms
4.1 Current AI Chat Interfaces
ChatGPT, Claude, Gemini, and their equivalents represent the dominant model of AI interaction today: conversational, turn-based, session-bounded, and instruction-centric. They are, in the framework's terminology, instruction-layer interfaces. The human formulates a specific request, the AI executes, and the result is returned in the same channel.
The Intent Layer diverges in several ways. First, chat interfaces are inherently session-based: each conversation starts from a limited context and ends when the tab closes. The Intent Layer is continuous: intent persists and accumulates across all interactions, building a deepening model of the human's direction. Second, chat interfaces are human-initiated: the AI responds but does not proactively surface observations. The Pulse is bidirectional: the system initiates signals based on its own analytical judgment. Third, chat interfaces are single-channel: everything flows through text in a rectangular window. The Intent Layer is multi-channel: signals reach the human through whatever sensory channel is most appropriate.
Chat interfaces are excellent for instruction-level interaction, and the Intent Layer does not replace them. At Depth Level 2, when the human is co-creating with an agent inside a specific artifact, the interaction may resemble conversation. But the framing is different: the conversation occurs within the context of the Field, the Pulse, and established intent, not as an isolated exchange.
4.2 Agent Orchestration Frameworks
Tools like LangChain, CrewAI, AutoGen, and similar frameworks approach the multi-agent problem from the infrastructure layer. They provide primitives for defining agent roles, routing messages between agents, and managing multi-step workflows. They are valuable engineering tools. They are not interaction frameworks.
The Intent Layer framework operates at a different level of abstraction. Where orchestration frameworks ask "how do we wire agents together?", the Intent Layer asks "how does a human direct and trust a system of agents?". These are complementary concerns. An orchestration framework might be used to implement the Agent Layer and parts of the Orchestration Layer in the Intent Layer architecture, but it would say nothing about the Sensory Stack, the Pulse, or the Field. The Intent Layer fills the gap between agent infrastructure and human experience.
4.3 Calm Technology
Amber Case's work on calm technology, building on Mark Weiser's foundational concept of ubiquitous computing, provides perhaps the closest existing precedent for several Intent Layer principles. Calm technology's core thesis, that the best technology is that which informs without demanding attention, is echoed directly in the Sensory Stack design, the ambient channel, and the default-OFF principle.
The Intent Layer builds on calm technology in two directions. First, it extends the concept from passive information display to active agent systems. Calm technology describes how information should reach humans. The Intent Layer describes how intent should flow bidirectionally between humans and autonomous systems. Second, it introduces the concept of depth levels, which calm technology does not address. Calm technology operates at what the Intent Layer would call Level 0: ambient awareness. The Intent Layer adds the continuous zoom from ambient awareness through focused engagement to deep creative flow, maintaining calm principles at every level.
4.4 Ubiquitous Computing
Mark Weiser's vision of ubiquitous computing, articulated in his seminal 1991 paper "The Computer for the 21st Century," imagined computation embedded invisibly into the environment. The Intent Layer shares Weiser's conviction that the most profound technologies are those that disappear. The default-OFF principle is a direct descendant of Weiser's observation that "the most profound technologies are those that disappear. They weave themselves into the fabric of everyday life until they are indistinguishable from it."
Where the Intent Layer departs from Weiser is in the nature of what disappears. Weiser imagined computational devices becoming invisible. The Intent Layer imagines computational agency becoming invisible: not just the hardware but the decision-making, execution, and coordination that currently consume human attention. The device is not the thing that needs to disappear. The work is.
4.5 Don Norman's Design Principles
Norman's principles of good design, particularly visibility, feedback, constraints, and mapping, apply to the Intent Layer but require reinterpretation. Visibility, in Norman's framework, means making relevant system state apparent. In the Intent Layer, this translates not to showing everything but to semantic compression: making exactly the right state apparent through exactly the right channel. Too much visibility is as harmful as too little.
Feedback in the Intent Layer is provided by the Cascade: when the human expresses intent, the system's reorganisation is itself the feedback, confirming that the intent was received and interpreted. Constraints are provided by the trust boundaries within which agents operate. And mapping, the relationship between controls and their effects, is replaced by something more fluid: the relationship between expressed intent and observed outcomes, which is learned and calibrated over time rather than fixed in a control layout.
4.6 Activity Theory and Distributed Cognition
Activity theory, particularly as developed by Engeström, offers a useful lens. The Intent Layer can be understood as redefining the division of labour in an activity system: the human shifts from subject-acting-on-object to subject-setting-direction-for-system-acting-on-object. The mediating artifacts shift from tools (which the human operates) to agents (which operate on the human's behalf).
Distributed cognition, as developed by Edwin Hutchins, is also relevant. The Intent Layer is fundamentally a distributed cognitive system: cognition is distributed across the human (intent, values, judgment), the orchestration layer (context, compression, routing), and the agent ecosystem (research, analysis, execution). The Pulse is the mechanism by which this distributed cognitive system maintains coherence, analogous to the shared representational spaces that Hutchins identified in his studies of naval navigation teams.
Open Questions and Gaps
Intellectual honesty demands acknowledging that the Intent Layer framework, as presented, raises as many questions as it answers. The following are the gaps I consider most significant.
5.1 Trust Calibration
The 5/95 principle requires extraordinary trust. The human must believe that the 95% of activity happening without their involvement is proceeding in alignment with their intent. How is this trust established? How is it calibrated over time? What mechanisms allow trust to deepen as the system demonstrates competence, and what triggers should cause trust to retract?
In my own experience, trust builds through a pattern I think of as "verify, then expand." Initially, the human verifies frequently, checking outputs, overriding decisions, maintaining tight control. As the system demonstrates consistent alignment with intent, verification becomes less frequent and the scope of delegation expands. But this process is idiosyncratic and fragile. A single significant misalignment can collapse trust that took weeks to build. A formal model of trust calibration in intent-based systems does not yet exist.
5.2 Multi-Human Intent
The framework as presented describes a single human directing a system. But most human activity occurs in social contexts: families, teams, organisations. What happens when multiple humans express intent into the same system? How does the system resolve conflicting intents? How does it handle hierarchical intent (a manager's direction vs. an individual contributor's preferences)?
The family context is particularly interesting. If two partners share a home system, whose intent governs the calendar, the household management, the children's activities? The technical challenges are significant but the social challenges are harder: intent systems will need to navigate the same interpersonal negotiations that humans currently handle through conversation, compromise, and occasionally argument.
Organisational contexts introduce further complexity. An Intent Layer for a company would need to propagate intent across hierarchical levels while preserving local autonomy. This begins to look like the challenge of organisational culture: how do you ensure that a CEO's strategic intent is faithfully interpreted four levels down without micromanagement?
5.3 Conflicting and Ambiguous Intent
Humans are not consistent. They express intents that conflict with each other and with their own prior statements. "I want to publish more frequently" and "I want every piece to be deeply researched" are both valid intents that create tension. How does the system handle this?
Current approaches, asking the human to resolve every conflict explicitly, do not scale. The system needs a model of intent priority, context-dependent weighting, and graceful degradation when contradictions cannot be resolved. It also needs to recognise when a contradiction is actually a creative tension that should be preserved rather than resolved.
5.4 Failure Modes
Several failure modes deserve analysis:
Intent misinterpretation. The system interprets "be more aggressive in outreach" as increasing email frequency when the human meant taking bolder positions in published content. Misinterpretation at the intent level has more severe consequences than misinterpretation at the instruction level because it affects an entire domain of activity, not a single task.
Compression loss. The semantic compression function filters out information that would have changed the human's behaviour. Because the human never sees this information, they cannot know it was lost. The system's very effectiveness at reducing cognitive load creates a blind spot.
Cascade amplification. An intent expressed casually ("we should probably think about entering that market") cascades through the system as if it were a firm strategic commitment, triggering resource allocation and agent spawning that the human did not intend at that scale.
Trust miscalibration. The human trusts the 95% too much or too little. Over-trust leads to drift: the system gradually deviates from intent without the human noticing. Under-trust leads to micromanagement: the human checks everything, defeating the purpose of the architecture.
5.5 Privacy and Agency Boundaries
The Intent Layer requires deep knowledge of the human: their calendar, their communications, their preferences, their patterns, their relationships. This creates significant privacy concerns. Who has access to the context model? How is it protected? What happens if it is compromised?
There is also a subtler concern about agency. As the system becomes better at anticipating intent, the boundary between "serving the human's intent" and "shaping the human's intent" becomes blurred. If the system consistently surfaces certain types of opportunities and suppresses others, it is not just executing intent but influencing it. The Intent Layer must grapple with the question of where facilitation ends and manipulation begins.
5.6 The Cold Start Problem
A new Intent Layer system knows nothing about the human. It has no intent model, no context, no pattern history. How does it begin? The cold start problem is not unique to this framework, but it is particularly acute because the value proposition depends on deep understanding that can only be built over time.
Possible approaches include bootstrapping from existing data (calendar, email, documents), explicit intent-setting sessions ("tell me your priorities for the next quarter"), and a gradual ramp from instruction-level interaction to intent-level delegation. But the experience during the cold start period will be significantly degraded compared to a mature system, creating a valley of disappointment that many users may not cross.
5.7 Economic and Institutional Questions
Who builds the Intent Layer? The framework as described requires integration across devices (phone, watch, desktop, ambient sensors), communication channels, data sources, and AI capabilities. No single company currently controls all of these touchpoints.
Is the Intent Layer a platform, a protocol, or a product? Each answer implies different economic structures, incentive alignments, and competitive dynamics. A platform risks the same attention-competition dynamics the framework is designed to avoid: if the platform operator's revenue depends on engagement, the incentive to be truly default-OFF is weak.
The compute economics are also significant. Running persistent agents, maintaining context models, and performing continuous semantic compression requires substantial computational resources. The cost of the Intent Layer cannot be trivial. Who bears it, and how is the value captured?
5.8 Accessibility
The multi-sensory Sensory Stack creates both opportunities and challenges for accessibility. For users with visual impairments, the shift toward voice and haptic channels may be beneficial. For users with hearing impairments, the visual Field may be primary. But the framework needs to account for users who cannot access one or more sensory channels, and ensure that no essential function is available only through a single modality.
Cognitive accessibility is another consideration. The framework assumes a human capable of abstract intent expression and comfortable with delegating control. For users with cognitive disabilities, or simply for users who prefer explicit control, the Intent Layer may need to support a spectrum from full intent-level interaction to more traditional instruction-level engagement.
5.9 Cultural Variation
The framework's emphasis on individual intent expression reflects a broadly Western, individualist orientation. In cultures where decision-making is more collective, where hierarchy is more pronounced, or where the relationship between a person and their tools carries different cultural weight, the Intent Layer may need significant adaptation. This is an area where the framework's theoretical foundations are thinnest.
Implications
6.1 For Software Design
If the Intent Layer framework is directionally correct, it implies that the current trajectory of software design, adding AI features to existing applications, is insufficient. The shift required is architectural, not incremental. Applications built around the assumption of continuous human operation cannot be retrofitted into an intent-based architecture any more than a horse-drawn carriage can be retrofitted into an automobile.
Software designers will need to invert their assumptions about attention. Instead of optimising for engagement (time on screen, interaction frequency, feature discovery), they will need to optimise for absence: how much value can the system deliver per unit of human attention consumed? The most successful intent-native software will be the software that is used the least, in the sense that it demands the least human involvement while delivering the most aligned outcomes.
This also implies the end of the application as the primary unit of software. If there are no apps, only depth, then the boundaries between what are currently separate applications dissolve. The calendar, the email client, the document editor, the project management tool: these become depth levels and workspace configurations within a unified Field, not separate products with separate logins, separate data models, and separate interaction patterns.
6.2 For Organisational Structure
Organisations are, in one sense, intent-propagation systems. A CEO establishes strategic intent. The organisation's structure, its layers of management, its processes, its culture, exists to decompose that intent into action across hundreds or thousands of people.
If AI agents can handle significant portions of this decomposition and execution, the organisational implications are substantial. Layers of management that exist primarily to relay and interpret intent may compress. The value of individual contributors shifts from execution capability to intent quality: the ability to express clear, coherent direction that agent systems can interpret and act on.
This does not necessarily mean smaller organisations. It may mean organisations with different shapes: flatter, with fewer middle layers, but with more humans focused on the Intent Layer functions of direction, values, taste, and judgment. The premium on human judgment may actually increase as the volume of agent-executed work grows, because the consequences of misaligned intent scale with the system's capability.
6.3 For Education
If intent expression becomes the primary human function in agent-augmented work, then education systems need to develop this capacity deliberately. Current education emphasises execution skills: how to write, how to calculate, how to code, how to analyse. These remain valuable but become less differentiating as agents become capable of performing them.
The skills that become critical in an Intent Layer world:
Articulating intent clearly. The ability to express direction, priorities, and values in ways that agent systems can interpret. This is closer to leadership communication than to technical instruction-writing.
Judgment under uncertainty. The ability to make the marginal decisions that the system surfaces: the novel situations, the ethical grey areas, the aesthetic choices. This requires wisdom, experience, and the capacity for nuanced reasoning that cannot be reduced to rules.
Systems thinking. Understanding how intent cascades through complex systems, anticipating second-order effects, recognising when a system's behaviour has drifted from intent. This is the capacity to steer a system you do not directly control.
Taste and aesthetic sensibility. The ability to say "not like that, like this" in ways that reflect genuine qualitative judgment. As more execution is automated, the human's editorial function, the capacity to recognise quality and direct toward it, becomes more, not less, valuable.
Trust calibration. The skill of knowing when to delegate and when to verify, when to expand the scope of autonomy and when to contract it. This is a meta-skill that governs the effectiveness of all other skills in the framework.
6.4 For the Nature of Work
The Intent Layer framework, taken to its logical conclusion, suggests a redefinition of work itself. If 95% of what we currently call "work" can be handled by agent systems, then "work" in the human sense becomes something quite different from what it has been for the past several centuries.
Work becomes the expression and refinement of intent. The creative acts that require human originality. The judgment calls that require human values. The relationship moments that require human presence. The reflective observations that require human meaning-making.
This is not a utopian claim. The transition will be difficult and uneven. Many forms of work will be displaced before new forms emerge. The economic and social structures built around execution-based work will strain. But the framework suggests a direction: toward a world where what makes work valuable is precisely what makes it human.
Conclusion
The Intent Layer is a framework, not a product. It describes a set of principles for designing human-agent interaction that respects both the exponential capabilities of AI systems and the finite, precious nature of human attention, judgment, and intent.
Its core claims are: the human's fundamental role in agent-augmented systems is intent expression, not task management; the interface should be absent by default and present only when human contribution is needed; interaction should flow through multiple sensory channels selected by context; the meeting point of human and AI intent, the Pulse, is a bidirectional space of multiplication rather than a unidirectional channel of instruction; and the application model should be replaced by continuous depth levels within a unified spatial environment.
Some of these claims will be refined or refuted as agent systems mature and the empirical evidence grows. The open questions identified in Section 5 are genuine gaps, not rhetorical gestures. This framework is offered as a starting point for a conversation that the technology industry has not yet had in sufficient depth: not "how do we add AI to existing interfaces" but "what interfaces does an AI-native world actually need?"
The answer, I believe, begins with recognising what is irreducibly human in a world of increasingly capable machines. Not the ability to execute. Not the capacity to process information. Not the speed of response. But the capacity to want something. To care about something. To decide what matters.
That is the Intent Layer. Everything else can be infinite. That part has to be you.
Fabio Oliveira is Head of Innovation at Kmart & Target Australia. He builds with AI agents daily and writes about the future of human-AI collaboration. This paper was developed from first-person experience designing and operating multi-agent systems, informed by ongoing conversations with both human colleagues and AI collaborators.