BDH-CQ combines in-context learning with reasoning outside the token stream

This article is part of our coverage of the latest in AI research.

Pathway researchers and collaborators have developed a reasoning model that can perform in-context learning without the need for generating intermediate thought tokens like LLMs.

The model, called BDH-CQ, combines in-context learning with recurrent latent reasoning, a trend that I’ve been following recently. The model updates its internal memory to internalize demonstrations presented at inference. It then reasons on the new problem iteratively in its continuous hidden state and only decodes the final answer.

In experiments on ARC-AGI-1, a benchmark of abstract visual reasoning tasks, a 150-million-parameter BDH-CQ model reached 29.5% (pass@2) at a computed inference cost of $0.0007 per task. According to the researchers, this puts it beyond the previously reported cost-accuracy Pareto frontier for the benchmark.

The result points to a different way of scaling reasoning. Today’s reasoning models often spend more compute by generating longer chains of thought. BDH-CQ spends that compute inside the model instead. If the approach transfers to language, math, and agentic coding tasks, it could provide another way to build AI systems that adapt to new tasks without paying for long reasoning traces on every request.

Reasoning outside the token stream

Current reasoning models usually work one token at a time. When prompted to reason through a problem, the model generates intermediate tokens that act as a scratchpad. Each step is projected into the model’s vocabulary, emitted, then processed again before the next step can be produced.

This approach has an important advantage: the intermediate reasoning is represented in the same language that humans and other software components can read. It also gives the model more computation by letting it extend its output.

But it ties computation to a serial channel. Every intermediate idea must be compressed into a token sequence before the model can continue. As reasoning traces grow, so do token consumption, inference compute, and latency.

Zuzanna Stamirowska, CEO of Pathway, told TechTalks that she sees a deeper limitation in using language itself as the medium for computation. “Chain-of-thought was an important breakthrough, but it treats language as the substrate for reasoning. We believe that reasoning and language are related but distinct phenomena,” she said. With text-based reasoning, she added, models have to verbalize intermediate computation as they navigate constraints, maintain alternatives, and converge on a strategy.

Latent reasoning takes a different approach. Instead of turning each intermediate state into tokens, the model repeatedly modifies its hidden representation and decodes an answer only when the computation is finished.

A continuous state can also carry information that is awkward to serialize into a single sequence. It can preserve partial hypotheses or candidate transformations at the same time. Another advantage of continuous reasoning is that it can let models explore multiple computational paths in parallel instead of spelling them out one after another.

Several architectures have already explored this direction. Coconut, for example, feeds a Transformer’s hidden state back into the model and progressively replaces verbal chain-of-thought steps with continuous representations during training. Other recurrent-depth and looped models repeatedly apply the same computation to a latent state to increase effective reasoning depth.

The missing piece is flexible task acquisition. Language models are good at learning a new task from a few examples in the prompt, often referred to as “in-context learning.” Many recurrent reasoning systems, including ARC solvers, instead adapt to evaluation tasks through optimization or task-specific representations.

BDH-CQ tries to combine these two capabilities: learn the rule from examples during inference, then reason about the new problem in latent space.

From BDH to BDH-CQ

BDH-CQ builds on Pathway’s Dragon Hatchling, or BDH, architecture. BDH differs from a conventional autoregressive Transformer in how it represents and moves information through the network. The researchers describe it as using “high-dimensional positive activations, low-rank communication, and a recurrent associative state.”

The model maintains a large internal feature space. Its activations are positive and ReLU-based, which means information can be represented as patterns of active features across that space. Instead of allowing every feature to communicate densely with every other one, BDH uses lower-dimensional transformations to move information between parts of the representation.

BDH also keeps a persistent recurrent state. Information from earlier inputs can remain in that state and change as new inputs arrive.

That creates a different computational pattern from the standard autoregressive Transformer. A Transformer reasoning model typically processes a token context, predicts another token, appends it to the sequence and repeats. BDH can continue changing its internal state through recurrence without requiring every computational step to become another generated token. BDH still includes linear attention, but recurrence and persistent state play a more central role in how computation proceeds.

The original BDH work focused on sequence modeling and language capabilities. The same layers have also been used in a Sudoku system that repeatedly refines a latent representation until it satisfies the puzzle constraints.

In Pathway’s previous Sudoku benchmark, BDH achieved a 97.4% solve rate across approximately 250,000 Sudoku Extreme puzzles. According to Stamirowska, reasoning models including OpenAI o3-mini and DeepSeek R1 scored 0% on the same benchmark. Sudoku is a useful example of the kind of workload recurrence targets: the model has to keep track of many interacting constraints and repeatedly refine a candidate solution before committing to an answer.

BDH-CQ extends that architecture with a system for learning transformations from demonstrations and applying them through recurrent latent reasoning. BDH-CQ adds an interface for learning visual transformations from examples and a structured latent workspace in which the model can reason over a query.

The mechanism separates two jobs that are often mixed together in a Transformer prompt: learning what the task is and working out the answer.

First, BDH-CQ processes demonstration input-output pairs one by one. Each example changes a recurrent contextual state. The model’s parameters do not change. Instead, the examples gradually build task-specific associations in memory.

Once the demonstrations have been processed, the query enters a separate latent workspace. The model repeatedly updates this workspace while conditioning on the contextual memory. After several iterations, it decodes the final state into an output.

The contextual state represents what the model has learned from the demonstrations. The latent workspace carries the computation needed to solve the current query.

Consider the ARC task below. The demonstrations contain two binary grids separated by a gray column. Their outputs reveal that the desired transformation is to mark cells occupied in both grids. Nothing tells the model that the concept is an intersection. It has to infer the relation from the examples and apply it to a new grid.

In BDH-CQ terms, the demonstrations shape the recurrent memory into a representation of that rule. The model then iterates over the new grid in latent space to produce the output.

There is still a limit to how deeply developers can inspect or reproduce this mechanism. The paper describes the system-level architecture but keeps the exact recurrent update rules, state dimensions, and parts of the training recipe proprietary.

BDH-CQ in action

BDH-CQ was mainly tested on ARC-AGI and the headline result shows the cost advantage. The 150M-parameter system solved 118 of 400 tasks under pass@2, or 29.5%, using around 0.85 H200 GPU-seconds per task. At the paper’s assumed price of $3 per H200-hour, that works out to $0.00070 per task.

Stamirowska attributes much of this efficiency to moving both reasoning and memory inside the model. “The dramatic efficiency boost comes from eliminating expensive token-based reasoning and keeping memory and reasoning inside the model rather than repeatedly reconstructing context,” she said. Instead of decoding each reasoning step into tokens and feeding those tokens back into the next step, BDH-CQ can continue computation in its recurrent state. 

The comparison is not perfectly apples-to-apples. Other systems in the paper’s cost chart use figures reported by the ARC leaderboard, which can represent API prices or hardware estimates. Still, BDH-CQ occupies a low-cost region of the chart that previous systems did not reach at comparable accuracy.

The more revealing experiments test what the model actually learns from demonstrations.

In one test, every task introduced a new color mapping. The model had to infer the mapping from examples and apply it to held-out sequences. BDH-CQ solved all 96 held-out outputs at rank one, including cases with eight simultaneous color bindings.

Other experiments tested whether the model could apply a demonstrated operation beyond the range shown in its examples. Boundary propagation remained correct on all 48 held-out outputs as the distance increased. Copying also remained at 48/48 as the number of target sites increased.

But the behavior changes sharply for other kinds of complexity. For example, when the model had to order objects by height, performance stayed near saturation through five objects, then dropped dramatically as the number of objects increased. 

The contrast matters for applications that depend on few-shot adaptation. Some failures can be addressed by giving the model examples that better cover the complexity it will encounter. Others remain execution bottlenecks even when the relevant pattern appears in context.

Composition is another weak point. The model solved rotation combined with relocation on all 72 held-out outputs. Reflection combined with relocation fell to 47/72. 

BDH-CQ also does not always apply an inferred rule consistently. On ConceptARC, it reached 77.92% pass@2 on individual test pairs but only 59.38% when a task counted as solved only if all three test inputs were correct.

What persistent state could change for agent engineering

The recurrent-memory approach could also have implications beyond standalone reasoning tasks. Much of today’s agent stack is designed around the fact that model calls do not automatically carry forward everything learned during previous interactions. Engineers retrieve conversation history, maintain external memory stores, summarize old context, compress prompts, orchestrate tools, and decide which pieces of information need to be injected again at each step.

“Every agent step can mean re-reading and re-summarizing context, while developers carry the operational burden of coordination, governance, reliability, and long-running workflows,” Stamirowska said. “The result is that latency and token usage grow alongside the complexity of the task, not because the underlying work necessarily requires it, but because the system has to repeatedly reconstruct what it already learned.”

A model that can preserve useful internal state could move part of this burden from the agent harness into the model. Instead of reconstructing the full relevant context for every call, an agent could carry forward task-specific associations or experience as state. Stamirowska describes the potential shift as moving “from context reconstruction to persistent intelligence.”

That would not eliminate the need for components such as RAG. Agents would still need to retrieve current documents, authoritative records, or information they have never seen before. But persistent model state could reduce the amount of infrastructure devoted specifically to reminding a model of information it has already processed.

The current BDH-CQ experiments do not yet demonstrate this in long-running enterprise agents. They show in-context adaptation within ARC-style tasks. Extending that behavior to persistent agent memory across tools, workflows, and longer periods remains a future application of the architecture.

Where BDH-CQ fits in recurrent reasoning

BDH-CQ also provides an interesting comparison with Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM), two recurrent architectures that have produced strong results on puzzles such as ARC, Sudoku and mazes.

HRM and TRM also perform repeated computation in latent state. But their ARC pipelines adapt to evaluation tasks through optimization. Demonstration pairs are augmented and used during training, and augmented puzzles receive learned task identities. A previously unseen task therefore requires backward-pass adaptation before evaluation.

BDH-CQ puts that adaptation into the inference process instead. Demonstrations change recurrent memory while the model weights stay fixed.

That distinction moves recurrent reasoning closer to how developers already work with general-purpose LLMs: provide a few examples of a bespoke task, then ask the model to perform it.

BDH-CQ also sits alongside another emerging approach: recurrent-depth or “looped” Transformers, which repeatedly run a shared Transformer block over an internal representation to add effective computational depth without adding new parameters at every step. A looped-Transformer technique has also been reported in connection with OpenAI’s GPT-6 Astra.

Stamirowska sees an architectural distinction between adding recurrence to a Transformer and designing recurrence into the model from the ground up. BDH was built around recurrent state from the start, while BDH-CQ further separates task memory from the latent workspace used to solve a query. The BDH-CQ paper similarly distinguishes its focus from recurrent-depth language models: those approaches mainly explore parameter-efficient depth and test-time compute, while BDH-CQ focuses on learning a new transformation from demonstrations and then executing it through latent recurrence.

This does not necessarily mean token-based reasoning disappears. A model can benefit from latent computation internally and still generate intermediate language when it needs to explain a decision, interact with a human, verify a step, or call an external tool.

The current study focuses on ARC-style visual reasoning. The researchers identify language, mathematical reasoning, harder ARC tasks and constraint satisfaction as future directions.

Recurrence also comes with historical baggage. Recurrent neural networks have traditionally raised concerns about difficult optimization, vanishing gradients, and sequential computation that is harder to parallelize than Transformer training. Pathway’s goal with BDH is to retain adaptive recurrent state without giving up the scaling and distributed-training techniques that helped make Transformers practical.

The early evidence is encouraging but not yet conclusive. The researchers report Transformer-like pretraining scaling behavior for BDH-CQ models ranging from 1 billion to 600 billion parameters. The architecture also inherits tensor-sharding patterns from BDH, which the researchers say make it suitable for distributed training and potentially for models at the trillion-parameter scale. Stamirowska said the team does not currently see “a fundamental scaling penalty that would make BDH-CQ impractical at large dimensions.”

Larger models and longer training runs will still be needed to determine whether the current limitations in extrapolation, composition, and consistent execution move with scale.

Developers will also have to wait before they can test the system directly. Pathway does not currently plan to open-source BDH-CQ’s model weights, full training code, or exact update and training recipes. Instead, the company plans to make the technology available as a service.

“Our roadmap is to expand from working with design partners to a research-preview API, followed by beta and then general availability,” Stamirowska said. This would give developers a way to experiment with the architecture without Pathway publishing the underlying implementation.

It also remains to be seen whether the same pattern works when the demonstrations are code, documents, tool traces or business rules instead of colored grids. If it does, latent reasoning could become one component of a broader AI stack in which models do more computation internally while exposing language when humans and software systems need it.

“I personally don’t see a future where latent agentic systems totally break from language for communication,” Stamirowska said. “I see latent reasoning and text-based reasoning complementing each other in production.”

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.