feat: harden unified memory and unify provider-aware thinking controls - #943
Open
alcholiclg wants to merge 42 commits into
Open
feat: harden unified memory and unify provider-aware thinking controls#943alcholiclg wants to merge 42 commits into
alcholiclg wants to merge 42 commits into
Conversation
…openai-compat transport
…x/runtime-robustness
The orchestrator now owns the write discipline around a backend: - schedule_add() runs the extraction-LLM + embedding cost (seconds) in a background task; flush_pending() is the teardown barrier so the last write is never dropped, and an inline fallback keeps writes when no loop is running. - retrieval/ingestion/flush serialize on one per-store asyncio lock (embedded qdrant underneath is lock-free single-client code). - a content-hash delta ledger (<base_dir>/ingest_state.json) makes each ingest send only messages the store has not seen; hashes are recorded only after a confirmed write, so a failed ingest retries naturally. - ingest_status reports the last outcome (state/count/error/pending) so a UI can show memory working instead of silence. Mem0Backend: per-turn retrieval cache (rounds 2..N of a tool-calling turn reuse round 1's search instead of paying an embedding round-trip each), on_messages returns the event count and propagates failures -- the orchestrator is the swallow-and-report layer now and needs the exception to keep failed messages un-marked for retry.
…-side close - add_memory(add_after_step) now fires only when a round closes the turn (assistant reply with no tool calls) and dispatches through the backend's schedule_add when available: tool rounds are intermediate state, and ingesting every round cost O(rounds x history) extraction calls where the closing ingest covers the whole turn. - an interrupted round advances the ingest ledger WITHOUT ingesting (mark_ingested): a half-finished answer is not durable conversational truth and must not be swept into the next turn's delta. - cleanup_tools drains scheduled ingestion (flush only -- memory instances are shared across agents of one store, so closing here would yank the store from a sibling agent); the new SharedMemoryManager.close_matching(base_dir) is the owner-of-last- resort that actually closes instances and releases the embedded store's exclusive file lock.
The number of recalled memories injected per turn was hardcoded twice (search default 20, then a [:10] formatting slice). MemoryConfig gains recall_top_k (default 10, read from the unified_memory node) and the mem0 adapter threads it through search and formatting — consumers can now size recall to their context budget.
…E) instead of scattered config fields, gated by personalization.enabled. When those files change mid-conversation the next user turn carries a durable <system-reminder> naming them, so the model can tell a changed file from its own faulty memory.
…end's MEMORY.md snapshot in step with edits made outside the agent. Also translates the memory tool descriptions and prompt headings to English.
…fig patch once instead of twice.
# Conflicts: # .gitignore # ms_agent/memory/unified/backends/mem0_adapter.py # ms_agent/memory/unified/orchestrator.py # setup.py
…ts last entry deleted, instead of leaving the previous round's block in place.
- an interrupt marks only its own round, and never messages a scheduled ingest still owns (both lost the write silently) - one shared instance per store, not per model, reconfigured in place - close() is terminal: a straggler can no longer reopen a released store - the store lock is per (loop, path), and search() takes it too - search() honours its limit; injected memories carry their date - memories are written in the language the user used
# Conflicts: # ms_agent/memory/unified/backends/mem0_adapter.py
…ters Thinking support is per-model with no naming rule, and an unsupported model may reject the whole request (DashScope returns 400) instead of ignoring the flag. So we ask, and on a refusal retry once with it off, remembering the model.
…t do different jobs
…orwards, and read OpenRouter's reasoning field
…f a smaller invented one
…the command the user approved
…a text file listing
…o a vision-disabled model neither claims nor disowns them
…x/runtime-robustness # Conflicts: # ms_agent/agent/llm_agent.py # ms_agent/memory/unified/backends/mem0_adapter.py # ms_agent/memory/unified/orchestrator.py
…tover credentials in the environment
wangxingjun778
approved these changes
Aug 19, 2026
…hat arrive mid-stream - images go out only when the model's own switch says so; a provider's declared vision capability no longer implies it - a 400 delivered on the first streamed chunk is repaired like an eager one - thinking refusals are repaired on the Anthropic and Responses paths too - a tool call the model is still writing is reported instead of nothing at all - an unreadable managed MCP config is logged instead of silently yielding none
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Change Summary
Unified memory reliability
recall_top_kandingest_intervalimmediately;close()terminal: drain scheduled writes, release the backend safely, and prevent delayed tasks or later reads from reopening a retired store.Provider-aware thinking controls
reasoning_effortsetting withauto,off,low,medium,high, andmaxlevels.reasoning_effortwiththinking_budget.enable_thinkingandreasoning_efforton DashScope because the switch and effort level control different behavior; translatemaxto DashScope’sxhigh.reasoning_contentand OpenRouter’sreasoningresponse field so proxied reasoning remains visible.MSA_DEBUG_THINKINGdiagnostics for inspecting the lowered request.Shell command permissions
curl,wget,ssh,scp,rsync,nc,netcat) instead of refusing them: as blacklist entries no mode or user answer could override them, and blocking by command name never stoppednpm,pip, orgitfrom reaching the network anyway — they are ask rules now, enforced on the ordinary permission path whereask_ruleshad been inert (allow_network: trueopts out).*indangerous_removal_pathsfrom being applied as an fnmatch glob, which marked every path dangerous and had the non-bypassable safety layer refuse everyrm,rm build/out.txtincluded;rm -rf /,rm *, andrm ~stay refused.allow_alwaysonls -lano longer releases the whole shell, and match a bare<cmd>against its own<cmd> *pattern.Multimodal image input
Image N: <filename>so later turns can reference them, and are rescaled or transcoded to fit each provider's limits before sending.file_system---read_fileinstead of base64 in the text channel — inline intool_resultfor Anthropic, hoisted into a following user message for the OpenAI family, whose Chat Completions schema restricts tool message content to text parts. Charge image blocks a flat per-image cost in the context estimator, which previously measured base64 length and could let a single image exceed the whole context budget.flatten_message_textin memory recall, compaction, hooks, session naming and history rendering so it cannot silently degrade, and add an optionalimage_readertool that describes an image through a separately configured vision model.Related issue number
Checklist
pre-commit installandpre-commit run --all-filesbefore git commit, and passed lint check.