Anthropic Gives Managed Agents Time to Learn From Past Work
The Dreaming research preview reviews completed sessions and writes memory updates intended to reduce repeated mistakes.
SAN FRANCISCO, Calif. - Anthropic has introduced Dreaming, a research-preview process that lets managed artificial intelligence agents review earlier sessions and update their working memory between tasks. This system is intended to help agents identify repeated errors, preserve successful approaches and improve over time without retraining the underlying model. By allowing agents to reflect on their own performance, the company is addressing a significant roadblock in the deployment of autonomous systems that must handle complex, multi-step workflows. This release follows a broader industry trend where developers are seeking ways to make large language models more useful in production environments without the immense computational costs associated with continuous fine-tuning or full model training cycles.
The introduction of Dreaming comes as the Silicon Valley artificial intelligence sector shifts its focus toward agency, the ability for models to not only generate text but to execute actions across software platforms. Anthropic, a public benefit corporation founded by former OpenAI executives, has positioned itself as a safety-first competitor in this space, often emphasizing reliability and interpretability. As organizations attempt to move past simple chatbots and toward agents that can manage data entry, customer support, or legal research, the stability of an agent's long-term memory has become a critical technical challenge that requires a more sophisticated approach than standard context windows.
During idle periods, the process examines past behavior, looks for patterns and produces memory notes that can guide future work. This 'offline' reflection period allows the system to analyze its own chain of thought and outcomes without slowing down the active response time during a live user interaction. Anthropic offers controls over how those updates are created and applied, ensuring that the resulting memory notes are structured in a way that remains useful for the agent’s specific objectives. The approach targets a common weakness in long-running agents: memories can become cluttered, contradictory or outdated as the number of completed sessions grows, eventually leading to performance degradation.
Industry analysts have noted that current AI models often suffer from a form of digital amnesia, where lessons learned in one session are lost by the start of the next unless they are manually hard-coded back into the prompt template. By automating this reflection with Dreaming, Anthropic is attempting to create a loop where the agent becomes progressively more specialized for the specific quirks of a company’s internal data and preferences. This mirrors the way human workers refine their professional habits through experience, but it applies that logic to the high-speed processing of machine learning systems.
Legal technology company Harvey reported a sixfold improvement in task completion after using the feature, according to coverage of the launch. For a sector like the legal industry, where precision is paramount and historical precedent informs every decision, the ability for an agent to learn from its previous document reviews or research errors is particularly valuable. This reported gain suggests that agents with better memory management can navigate complex professional domains with far less human intervention than previously thought possible.
However, while the results from Harvey are notable, it is important to remember that they come from an early design partner and may not generalize across industries. Performance will depend on task structure, evaluation quality and the accuracy of the memory selected for future use. Different sectors have varying tolerance levels for error, and a memory update that works well for a creative task might be disastrous if applied to a strictly regulated field like healthcare or finance. The variability of real-world data remains a primary hurdle for any system attempting to self-optimize based on historical performance.
Dreaming shifts some improvement work from model training into the application layer, change that reflects a maturing AI ecosystem. By moving the learning process closer to the user’s specific data rather than the base model weights, Anthropic is providing a more targeted form of personalization. This can make an agent more responsive to a specific organization, but it creates a new governance problem. When an agent is permitted to write its own memory updates, the transparency of why it chose a certain action may become obscured by layers of previous, automated reflections.
The risk in this shift is that an incorrect lesson written into memory may influence many later decisions. If an agent misinterprets a user’s correction or draws the wrong conclusion from a failed task, that error could become embedded in its working memory as a 'best practice.' Over time, an agent could optimize for a narrow metric while moving away from the user's actual intent. This phenomenon of drift is well-documented in machine learning, and it becomes even more difficult to manage when the model is autonomously updating its own operational logic between sessions.
To mitigate these risks, organizations testing the feature should retain version history, human approval for consequential changes and evaluation sets that detect regressions. Without these safeguards, the benefits of improved efficiency could be offset by the introduction of hidden biases or systemic errors. Governance frameworks will likely need to evolve to include 'memory auditing,' where human supervisors periodically review the notes an agent has written for itself to ensure they align with organizational policy and safety standards.
The competitive landscape for AI agents is intensifying, with Google, Microsoft, and OpenAI all pursuing different methods for persistent memory. Some have opted for large vector databases that store every past interaction, while Anthropic’s Dreaming suggests a more curated, distilled approach. This strategy of identifying patterns and extracting core lessons, rather than simply recording a transcript of every action, could prove more efficient in terms of compute and more effective at preventing the 'memory clutter' that typically bogs down long-term agent performance.
The concept is promising because production agents need to adapt after launch. In a fast-moving business environment, a static model will quickly become obsolete as software interfaces change and company workflows evolve. A system that can reflect and learn provides a level of durability that is required for enterprise-grade tools. However, the successful implementation of such a system is not guaranteed. Its success will depend on whether Dreaming produces useful, reversible learning instead of accumulating confident but flawed habits.
As Anthropic moves this feature from a research preview toward a more general release, the focus will likely remain on its stability in high-stakes environments. The ability to reverse a memory update is as important as the ability to create one, as it ensures that any 'hallucinated' lessons can be purged before they cause operational damage. For now, the developer community is watching to see if this shift toward reflective agents will set a new standard for how AI interacts with long-term data and user feedback.
Ultimately, the goal of the Dreaming preview is to bridge the gap between fixed-state models and truly adaptive digital colleagues. If the technology can consistently replicate the successes seen with partners like Harvey, it may represent a significant step toward agents that are not only smarter but more reliable in the face of repetition. The coming months will be a test of whether an AI can truly be trusted to judge its own past performance, or if the human in the loop remains the only reliable source of correction in the workplace.
Sources
Written by
The Company Wire Staff
Reporting from The Company Wire newsroom. Staff bylines cover funding rounds, product launches and company news verified against primary sources.



