I post-trained Qwen3 30B-A3B into a text interaction model. It follows typing as a timestamped stream and chooses one action at each step: stay idle, mark a span, delegate a lookup, integrate a pending result, skip stale information, or respond. Most of the time, it stays quiet.
The project started with a practical question about Thinking Machines' interaction models: what would the same idea look like in text, where the model sees partial drafts but should rarely intervene?
The resulting model can mark relevant words, look up information while a sentence is still being written, hold onto the result, and incorporate it later if it remains useful. It can also discard information after the user has moved on. Getting these actions to work required relatively little training; getting the model to stop volunteering correct but unwanted information turned out to be the more interesting problem.
Why text?
The interaction models released by Thinking Machines are trained to operate on continuous streams, which lets them handle interruptions, overlapping speech, and ongoing changes in context without relying entirely on an external turn-taking system. Their demonstrations make a strong case for this approach in audio and video, but I was curious whether the same general idea could be useful in a text interface.
The version I wanted behaves more like someone quietly following along while you work. It can notice that you are missing a fact, recognize that you asked it to flag a certain kind of word, or determine that information it found a moment ago is no longer relevant. Most of the time, however, there is nothing useful to add, and the correct action is to do nothing.
Demos
A good demo needs to show more than whether the model can perform an action. The behavior has to make sense in context, and the model should avoid doing the same thing in situations where it would be distracting. Otherwise, it is difficult to tell whether the system understands the interaction or simply takes every available opportunity to act.
The held fact
The first example is based on a note written after a Spurs-Thunder Game 7. While drafting a recap, the writer mentions that they cannot remember the final score. The model notices the missing information and starts a lookup before the sentence is finished.
The result comes back as Spurs 111, Thunder 103, but the model does not immediately insert it into the draft. It keeps the information available while the user continues typing, then adds it when the writing reaches a place where the score naturally belongs.
The model starts a lookup mid-sentence, holds the result, and waits for a natural opening.
The timing matters because the behavior would be much less interesting if the user had explicitly stopped writing to ask for the score. It also would not help if the model inserted the result into an unrelated sentence simply because the lookup had completed. The point of the example is that the model can notice an unresolved question, let another process handle the search, and preserve the result until the draft gives it a reasonable place to use it.
Marks
The model can also follow instructions about what to notice while the user is typing. If asked to flag animals, it marks examples as they appear in the text. If asked to identify filler words, it underlines words or phrases such as “um” or “you know” without taking over the conversation.
The animal examples include words such as quokka, kestrel, and wombat, even though those particular species were not part of the training examples. The base model already knows what an animal is, so the main thing it needs to learn is how to apply an active instruction to a live stream and point to the right span.
These marks appear as underlines or small annotations rather than full responses. The model should also stop when the instruction is withdrawn, which matters because an interaction model that keeps following an outdated instruction is often more irritating than one that misses an example.
Doing nothing
The least visually impressive demo is probably the most important one. The user types normally, pauses, revises a sentence, and continues writing while the model repeatedly chooses idle.
Nearly half of the training decisions are idle examples, and I kept those examples at full weight. Downweighting them produces a model that becomes restless and starts looking for excuses to intervene, even when the user has not asked for anything and there is no pending information to surface.
For an interface that is watching continuously, the absence of unnecessary behavior is part of the capability. A model that always finds something to say may look more active in a recording, but it is considerably less pleasant to have open while you are trying to write.
Representing the interaction as a stream
The model receives a serialized history containing events from the user and any background tools. At each decision point, it predicts a single action.
A simplified example looks like this:
<stream_event index=14 <t+650ms>
source=user
state=active
revised=false>
trying to get this game 7 recap into shape and
i'm blanking on the final score from spurs okc
</stream_event>
<stream_event index=15 <t+1800ms>
source=tool
state=paused>
{"result": "San Antonio beat Oklahoma City 111-103 in Game 7"}
</stream_event>
<PREDICT_THIS_ACTION>
User typing enters the stream as snapshots of the current text, while tool results enter through the same queue. The model is instructed to treat the stream as its complete context, which keeps the interaction grounded in what has actually happened rather than in hidden state supplied elsewhere.
The action space contains six possibilities:
idle
mark(kind, span, style)
delegate(tool, args)
integrate(text, source)
skip(reason)
respond(text)
Only respond claims the conversational floor. The other actions either update the interface, start background work, or decide that no visible intervention is appropriate. This separation is what allows the system to underline a word or hold a retrieved fact without behaving as though the user has handed control of the interaction to the model.
Each synthetic event is paired with a target action, so one stream can produce many training examples. The basic unit of supervision is a decision about what to do at a particular moment, rather than a complete assistant response.
Building the training data
I generated interaction scenarios by splitting text into short chunks of roughly two to six words and adding realistic pauses, revisions, and occasional conversational backchannels. The scenarios cover several behaviors, including factual lookups, instruction-conditioned marks, and situations where an instruction appears inside quoted text and should therefore be ignored.
The scaffolding that generates these scenarios also produces provisional actions, but those actions are not used as training labels. They exist only to create the structure of the interaction. A stronger model relabels each decision point according to the intended behavior, including whether an instruction is active, whether a marked span is complete, whether a missing fact should trigger a lookup, and whether a returned result is still relevant.
That separation is important because the heuristics are useful for generating examples but are not reliable enough to define the behavior I want the model to learn. If I trained directly on them, the student would inherit the scaffolding's mistakes and arbitrary assumptions.
The final interaction dataset contains about two thousand teacher-labeled decision points. I mixed those examples with general assistant data at approximately a two-to-one ratio to preserve the underlying model's ordinary assistant behavior.
I also reviewed every batch manually. Automated checks can confirm that an action has the correct format, but they cannot reliably tell whether the action would feel inappropriate while someone is actively writing. Several examples passed all of the structural checks and still made little sense when read as an actual interaction.
Training
The first stage uses supervised fine-tuning on the labeled streams, with a LoRA adapter trained through Tinker on top of Qwen3 30B-A3B.
This is enough to teach the model the mechanics of the action space. It learns to recognize when a lookup might be useful, how to emit a mark, and how to incorporate information returned by a tool. It also learns a less desirable habit: surfacing retrieved information even when the moment for using it has already passed.
At first, this looked like an ordinary data-quality problem. The more interesting issue was that the teacher itself behaved differently depending on how it was asked to make the decision.
When asked to generate the correct action for a stale factual lookup, the teacher tended to integrate the result anyway. When shown two candidate actions, one that inserted the stale fact and one that skipped it, the same teacher reliably preferred the skip.
In other words, the teacher could recognize the better behavior without consistently producing it on its own. Straightforward imitation transfers the generated behavior, so it also transfers the teacher's tendency to be overly eager.
The problem becomes clearer if you try to score the actions using factual correctness alone. The retrieved score is accurate, and the marked animal really is an animal. In the appropriateness examples I generated, a correctness-based reward preferred the intrusive action every time, even when that action no longer made sense for the user.
The missing information concerns whether the action is wanted in the current context. A fact can be correct while still arriving too late, and a valid annotation can be unwelcome when the instruction that justified it is no longer active.
To address this, I added a short direct preference optimization stage after supervised fine-tuning. The preference pairs compare a correct but unwanted action against a more appropriate alternative, such as skipping a fact that the user has already supplied. I mixed in a small amount of low-learning-rate supervised replay so the model would retain the behaviors learned during the first stage.
The preference run only takes tens of optimization steps and a few minutes of wall-clock time. Its purpose is not to add a new capability; the supervised model already knows how to produce the relevant actions. It changes which action the model prefers when several technically valid options are available.
Why I did not use on-policy distillation
Tinker's on-policy distillation recipe would normally be an appealing way to transfer behavior from a stronger teacher into a smaller model. The student generates its own outputs, and the teacher provides token-level information that encourages the student to behave more like the teacher on the states it actually visits.
For this project, the problem is that the teacher's generated behavior already contains the mistake I am trying to remove. It tends to integrate information after the relevant moment has passed, even though it can identify a better action when explicitly asked to compare alternatives. Distilling its token preferences would therefore strengthen the same over-integration pattern that the preference stage is meant to correct.
The structure of the action space also makes token-level distillation less informative than it would be for a long response. Each output is short and constrained by a fixed action grammar, so most of the tokens are already determined by the format. The important decision is which action to choose, especially in situations where the alternatives are idle, mark, integrate, or skip.
A token-level objective spends much of its attention on tokens that do not distinguish the behavior I care about. Preference pairs focus directly on the decision.
I still think on-policy distillation is useful when the main problem is transferring a capability from a stronger model into a smaller one, and it may become relevant again for improving delegation behavior. Here, however, the supervised model already has the ability to skip. The issue is that it does not choose to skip often enough.
Engineering the runtime
The runtime is built around a turn-based model, even though the interface needs to feel as though the model is following the user continuously. The model chooses the action, while the surrounding system handles sampling, execution, and a small number of constraints needed to keep the interaction usable.
Sampling the text
The browser samples the text area approximately 350 milliseconds after the first keystroke, every 650 milliseconds while typing continues, and every 1.8 seconds after the user pauses.
Each request sends the complete current text rather than a set of incremental edits. The server appends that snapshot to the interaction history as a timestamped event and rebuilds the serialized stream for the next model decision.
Full snapshots make revisions easier to handle because the latest version of the text is always available directly. They also reduce the amount of bookkeeping needed when requests arrive late or are dropped.
Inference runs serially. If a new snapshot arrives while the model is still processing the previous one, it replaces the contents of a single pending slot. Under heavier load, the model sees a coarser sequence of snapshots, but the user can keep typing without waiting for inference to finish.
Keeping the model from taking the floor
When the user is actively writing, the runtime suppresses full conversational responses. Marks, fact cards, and integrated text appear as annotations instead, which allows the model to provide information without treating every action as a new assistant turn.
There is also a thin execution layer around the policy. If the user asks the model to identify filler words, the model cannot start marking animals simply because it notices them. If an apparent instruction appears only inside a quotation, the runtime treats it as inactive. Actions rejected by these checks are recorded in an audit trail rather than silently discarded.
The visible interface shows only the actions that were actually executed, while the audit trail records everything the policy attempted. I used that trail to verify the demo recordings and to distinguish model errors from problems introduced by my own runtime.
In early versions, a surprising amount of the awkward behavior came from recovery heuristics rather than from the trained policy itself. At normal typing speed, the model often produced cleaner spans than the surrounding scaffolding allowed it to display.
Reducing latency
The earliest hosted LoRA setup was far too slow for the interaction I wanted, with some requests taking minutes per decision. I merged the adapter into the base model, quantized the result to four bits for local inference, and added a prefix key-value cache.
Successive snapshots share most of their prompt, so the runtime can reuse the longest common token prefix and prefill only the new portion. With those changes, per-tick latency dropped from roughly fifteen seconds to roughly one second.
This still does not reproduce the architecture of a genuinely streaming interaction model, where state persists naturally and new input extends the existing cache. It does make a conventional turn-based model fast enough to approximate the interaction in a text editor.
What still breaks
The system can still hallucinate while waiting for a lookup. If a tool result has not arrived, the model may guess the missing score and try to insert it anyway. The runtime now blocks integrations before lookup_ready and rejects text that is not supported by the actual returned result. Those blocked actions remain visible in the audit trail.
The model also sometimes starts a second lookup while the first one is still in progress. One of the existing demo recordings includes duplicate fact cards for this reason. I think this belongs in the same category as the over-integration problem: the model can issue a lookup, but it has not learned sufficiently reliable preferences about when another lookup would be redundant.
Skipping introduces its own execution issue. If the user remembers the score and types it before the model surfaces the retrieved result, the correct behavior is to discard the pending integration. The model can emit skip, and the runtime also converts an exact duplicate integration into a skip so that a queued action does not repeat text the user has already written.
Another limitation comes from the event schema. Several people instinctively tried to mark spans for the model, but the current representation only allows assistant-generated marks. Supporting user-created marks would require changes to the event types and training data rather than another pass over the existing examples.
The larger architectural limitation is that the system has no independent clock. Every action is triggered by an event entering the queue, so the model cannot decide to do something simply because five seconds have elapsed. Recurring reminders would require timer events, an additional action that can gently surface information without taking the conversational floor, and training examples for situations that the current dataset does not contain.
Even with those changes, this runtime would still operate on a scale of seconds rather than the much shorter intervals available to models designed natively for continuous interaction.
The next improvements are mostly about better preference data: avoiding duplicate delegation, improving the timing of integrations, and making skip behavior consistent enough to appear reliably outside carefully selected demos. There is also runtime scaffolding left over from earlier versions that the model may no longer need, and I would rather remove it than keep adding rules around behavior the policy should eventually learn.
What I found most useful about the project was that the underlying model did not need to learn much new factual knowledge. It already knew what animals were, how to interpret a retrieved score, and how to follow ordinary instructions. The harder part was learning how those capabilities should behave when the user is still in the middle of doing something.
Citation
Please cite this work as:
Rajan Agarwal, "Tiny Interaction Models", rajan.sh, Jun 2026. https://www.rajan.sh/tiny-interaction-modelsOr use the BibTeX citation:
@misc{agarwal2026tinyinteractionmodels,
author = {Rajan Agarwal},
title = {Tiny Interaction Models},
year = {2026},
howpublished = {rajan.sh},
note = {https://www.rajan.sh/tiny-interaction-models},
}