UpsQuad agent memory assessment
Independent source, research and development-runtime assessment, 20 September 2026.
This is a research and implementation assessment, not an approved PRD or delivery commitment. Runtime observations describe the development environment on the audit date; tenant production behavior was not verified.
Assessment
The seven proposed categories are useful for designing a persistent enterprise agent, but they are overlapping dimensions rather than seven independent components. Working memory is active state; semantic, episodic and procedural describe remembered content or capabilities; external and parametric describe storage; retrieval is an operation; prospective memory concerns acting on future intentions. A single runbook can be external semantic knowledge and procedural guidance. Efficient stateless tasks do not require all seven. A persistent SRE teammate benefits substantially from all these functions, but tenant-specific model training is optional.
UpsQuad has substantial context management, durable memory storage, governed extraction and retrieval, and durable workflow machinery. It has not yet demonstrated a reliable, continuously improving agent across all supported runtimes and tenant deployments. Episodic case retrieval, relevant procedural reuse, generalized future commitments, and measured downstream benefit are the largest capability gaps.
Method and limits
- Core source pinned to
616fe390e03cc1d743372b470b3c2e6ef48990f1; frontend source-management evidence pinned to clientbf2fffdee159867131e780cfa933f277167b4195. - Read implementation, composition roots, tests, deployment manifests, memory architecture and product requirements; checked live GitHub issue states.
- Inspected allowlisted non-secret flags in running dev containers; queried aggregate counts in a read-only database transaction at approximately 2026-09-20 19:30 UTC. No memory contents, user identifiers or transcripts were collected.
- Executed scoped offline tests. No live model evaluation, end-to-end tenant test, production/GCP inspection, feature activation or runtime mutation.
- Deployed revisions differ from audited main: Context Engine
b96f6ed9, workerd30b3209, orchestrator/memory MCP/Model Gatewayb3abef4d. Therefore code findings and runtime observations are separate evidence classes. - Maturity labels below describe implementation scope, not measured completion percentages. No agreed acceptance checklist establishes a defensible percentage, and averaging overlapping categories would double-count work.
Research and implications
CoALA distinguishes working, episodic, semantic and procedural memory and treats retrieval as moving long-term knowledge into working state. Working memory can persist outside the model context and supply selected input to each call. This supports the architecture here; it does not prescribe seven mandatory services. Cognitive Architectures for Language Agents
RAG explicitly combines model-parameter knowledge with an external index. External document storage and vector embeddings do not mean the agent model's weights have learned tenant facts. For changing enterprise knowledge, explicit retrieval also makes source attribution and correction practical. Retrieval-Augmented Generation
Lost in the Middle found that relevant information's position affected performance in the models tested. It motivates testing selection and compaction, rather than equating a large context limit with reliable memory. It is not a measurement of every current model. Paper
Voyager demonstrates an executable skill library with iterative execution feedback without model fine-tuning, in Minecraft. The transferable design idea is validated procedural reuse; its results do not establish enterprise-agent performance. Paper
LongMemEval evaluates extraction, reasoning across sessions, temporal reasoning, updated knowledge and abstention. These are useful dimensions for an UpsQuad evaluation suite beyond retrieving a matching document. Paper
PM-Bench evaluates remembering and executing intentions at future cues while other work continues. It separates remembering an intention from acting at the right moment. A newer September 2026 preprint explores a typed intention store with lifecycle logic in code. These support evaluating explicit future-intention state and trigger handling; neither is proof that UpsQuad's scheduler delivers this capability. PM-Bench, Typed intention stores
Coverage by category
| Category | Meaning and SRE example | Implemented coverage | Important unfinished work |
|---|---|---|---|
| Working / in-context | Current objective, constraints, observations and plan: “Investigate this CI failure without restarting production.” | Substantial. Context assembly, selective token packing, protected policy layers, compaction, recent history, checkpointed plans and tool state. Portal chat and LangGraph consume Context Engine assembly. | Prove fact retention under real-model compaction and restart continuity for every executor. Claude SDK checkpoint stores configuration; the SDK owns conversation state, so full recovery cannot be inferred from that checkpoint. |
| Semantic | Durable facts/preferences: “Service A deploys in region X.” | Substantial implementation; tenant rollout incomplete. Typed learned facts/preferences, extraction, embeddings, active-only semantic recall, provenance, review, correction and forgetting. Shared knowledge sources provide another factual store. | Current-model quality evidence, conflict/expiry policy, operational rollout and team-memory ownership. Agent-memory recall is org+agent scoped; a team_preference label does not make a row shared team memory. |
| Episodic | A past incident: “Last Tuesday this symptom came from exhausted connections; this fix succeeded.” | Partial. Ordered session/tool events, timestamps, workflow attempts, transcript APIs, and extraction from completed sessions. | No first-class analogous-incident retrieval was found in audited paths. Add situation/action/outcome summaries, temporal and cross-session search, and evidence links. A transcript archive alone does not establish useful episodic recall. |
| Procedural | Reusable know-how: “For symptom X, check Y before trying Z.” | Partial to substantial. Procedures extracted as mistake/fix/trigger records, reviewed persistent lessons, warm-start and PLAN injection, reusable workflow definitions. | PLAN recall ignores semantic relevance to the current goal. No demonstrated loop promoting outcome-validated lessons into versioned executable skills or proving fewer repeated errors. |
| Retrieval / external | Accessing knowledge outside active context: runbooks, documents, tools and stored memories. | Substantial. Vector plus lexical/recency retrieval, rank fusion, document registry/full reads, web/Drive ingestion, scoped source access, periodic refresh and knowledge MCP tools. | Unified query-time upstream freshness/access parity and connector coverage. Gateway routing does not automatically inject Context Engine memory. External agents need actual tool use or adapter integration. |
| Parametric | Knowledge and skills in model weights: general Linux or language competence. | Provided by the selected model. Model serving, catalog and endpoint governance are present. | No platform-managed tenant fine-tuning/adapter-learning pipeline found. This is optional, not a prerequisite for memory-capable MVP. Customer-provided trained endpoints remain possible. |
| Prospective | Remembering to act later: “When this deployment is healthy, notify the owner; stop if cancelled.” | Partial. Checkpointed todos, Temporal schedules, durable timers, event waits, governed polling and timeout paths. | No generalized conversational commitment capture/lifecycle found: owner, trigger, permission, cancellation, expiry, completion evidence and escalation need a coherent product model. Tenant manifests currently disable Temporal. |
Architecture map
Solid edges show implemented paths in the audited architecture, subject to runtime and deployment gates. Dashed edges show absent or incomplete connections. This is a functional map; boxes need not be separate services.
All paths also need authorization, provenance, correction/forgetting, budgets and quality measurement. Memory content must retain its source and trust boundary when used as model context.
Development-runtime observations
| Observation | Measured result | Interpretation |
|---|---|---|
| Extraction and assembly recall | Both enabled in running dev Context Engine | Flags establish configuration, not successful current model calls. |
| Warm-start / PLAN memory flag | Database override memory.warmstart_enabled=true | Enabled in dev despite false compile default. |
| Near-duplicate collapse | Disabled | Do not claim active broad consolidation. |
| Temporal | Enabled in dev Context Engine and orchestrator | No scheduled workflow was executed during this audit. |
| Stored memories | 99 total; 46 active, 53 rejected | Real persisted rows exist. Count is not a quality metric or proof of tenant production usage. |
| Active categories | 16 learned facts, 14 procedures, 15 work artifacts, 1 team preference | Current corpus is small; no active conversation rows in this table. Session events are separate. |
| Extraction run records | 207 succeeded; 9 dead-lettered | Historical activity; a succeeded run can produce zero candidates. |
| Latest extraction update | 2026-08-31 17:40 UTC | No updates in the preceding seven days. |
| Session lifecycle correlation | Latest termination also August 31; zero recent terminated/crashed sessions with transcripts | Extraction inactivity alone does not demonstrate a broken consumer. |
| Captured events | 1,267 total; 25 in preceding seven days | Recent event recording does not imply a terminal-session extraction should already occur. |
| Human review markers | 15: nine confirmed, six rejected | Selected review outcomes; not a random precision estimate. |
| Sampled human-audit verdicts | Zero recorded on memory rows | Confirms the gap tracked in #2654 in this dev dataset. |
The tenant compose file defaults extraction and semantic assembly recall off, while the media.net overlay explicitly pins them off. They also disable Temporal. This is checked-in target configuration, not an inspection of the running tenant installation. The memory feature's delivered issue status therefore does not imply that it is active for that customer.
Findings that affect readiness
- The complete experience depends on the runtime. Portal chat and LangGraph call Context Engine assembly. Claude SDK has a different state owner and can receive warm-start lessons. External agents get explicit MCP access. Model Gateway forwards request bodies; using that gateway does not prove automatic memory retrieval or injection.
- Relevant lessons are not guaranteed at planning time. The PLAN recaller accepts the goal but selects active procedural/factual rows by class and composite importance/recency/recall score. Semantic assembly recall is stronger than this planning path. Converge the two before claiming situation-specific learning.
- Human quality feedback remains incomplete. #2654 is open, and live dev recorded zero sampled-audit verdicts. The frozen evaluation floors an old, non-attributable model stamp at 8/11 confirmed; it explicitly has no human precision floor for attributable stamps. Neither that legacy 72.7% nor the current 9/15 selected review decisions estimates current general extraction quality.
- Production extraction billing is an explicit blocker. #2669 remains open: extraction lacks required usage attribution and budget enforcement. Close the metering/enforcement acceptance criteria before enabling this path for the gated offering.
- Consolidation and forgetting have different levels of completion. Exact deduplication and corrective supersession exist. Near-duplicate retirement is disabled in dev; broad synthesis/contradiction resolution is not demonstrated. Org retention exists, but loop-written rows omit per-row expiry, so per-type decay claims are not fulfilled.
- Retrieval quality evidence is narrow. The recorded benchmark has 40 synthetic documents and 20 short single-hop queries. Nomic's 20/20 primary recall@5 is a useful fixture result, not production accuracy. It is vector-only and does not test end-to-end memory usefulness, multi-hop episodes, unanswerable questions or tenant permission changes.
- Enterprise knowledge governance is unfinished. Last-sync document reads explicitly report unverified freshness; standalone live Wiki.js tools are a distinct available path. Team Knowledge #2995 remains open. #3259 also tracks unresolved dev database-role isolation, with production explicitly unverified; it prevents treating existing authorization code as proof of all deployed isolation. This audit did not test an exploit or inspect production.
Recommended deliverables and acceptance evidence
| Priority | Deliverable | Completion evidence |
|---|---|---|
| P0 | Prove the full memory loop on each supported deployment/runtime | An authorized session creates a valid memory; a later relevant task retrieves it; a correction supersedes it; a forgotten/rejected memory is absent; permissions and restart behavior hold. Trace source to use without exposing secrets. |
| P0 | Close rollout controls: extraction usage/budget enforcement, human audit decisions, detector/model readiness and monitoring | #2669 acceptance passes; sampled reviews receive human verdicts and stalled review alerts fire; tested feature enable/rollback plan per offering. |
| P0 | Establish a representative memory evaluation suite | Compare memory-on vs memory-off with fixed tasks/models. Measure task success, repeated errors, retrieval accuracy, unsupported answers, latency, cost, and memory-write precision. Include updated facts, ambiguous/no-answer queries and authorization changes. Agree thresholds before declaring completion. |
| P1 | Make procedural recall goal-conditioned | Relevant procedure ranks above irrelevant high-importance lessons; misleading/obsolete lessons are excluded; repeated incidents improve against the memory-off baseline. |
| P1 | Add episodic summaries and authorized cross-session case search | Agent finds an analogous incident, cites its actions/outcome/time, distinguishes a failed remedy, and answers temporal questions across sessions. |
| P1 | Add a future-intention model backed by the existing durable workflow engine | A conversational commitment becomes a reviewed/authorized intention; time and event cues trigger it; cancel/change/expiry work; restarts and duplicate events do not cause duplicate side effects; completion has evidence. |
| P1 | Finish unified knowledge freshness and upstream access parity | A changed or revoked source is handled correctly at query time; output states freshness; external-agent integration tests prove retrieved context actually reaches reasoning. |
| P2 | Consolidation, decay and validated reusable skills | Contradictions preserve provenance; obsolete memories retire safely; approved procedures are versioned and evaluated before reuse/promotion. |
| Optional | Tenant-specific parametric learning | Consider only after evaluations identify a behavior gap retrieval/procedures cannot adequately solve; require training data governance, model evaluation and rollback. |
Suggested SRE acceptance scenario: investigate a failed deployment, retain one verified incident episode and one useful lesson, restart the worker, retrieve the relevant case next week, update a changed environment fact, and carry out a separately authorized follow-up only when its health condition is met. Test a cancellation and revoked source access too. This connects the categories into observable agent behavior.
Verification receipts
All selected offline tests passed on the audited source:
GOMAXPROCS=2 go test -p 1 ./internal/context/memory/extraction/eval ./internal/context/memory/ranking
GOMAXPROCS=2 go test -p 1 ./internal/context/memory/extraction ./internal/runtime/transcript
GOMAXPROCS=2 go test -p 1 ./internal/context/assembly ./internal/runtime/autoloop -run 'TestScenario4_CompactionLongRunSurvival_AsyncLeadsNoTruncation|TestInjectPlanMemories|TestRenderPlanMemoryBlock' -count=1
GOMAXPROCS=2 go test -p 1 ./internal/workflow ./internal/temporalwf/interpreter -run 'TestSetWorkflowEnabled_ScheduleLifecycle|TestSetWorkflowEnabled_NoReconciler_ConfigOnly|TestWaitPoll_ExternalCIObservation_(ResumesOnSuccess|NeverGreenRoutesOnTimeout|UnknownConclusionIsNotSuccess)' -count=1
These validate selected contracts with fixtures/fakes. They do not establish real-model retention, production availability or improvements in task outcomes. Retrieval/isolation integration tests were inspected, not executed.
Implementation evidence
All core links below pin the audited revision.
- Context assembly pipeline, budget packing, LangGraph integration, SDK checkpoint limitation.
- Extraction prompt, semantic memory recall, review lifecycle.
- Session event schema, bounded same-session history, PLAN relevance limitation.
- Hybrid retrieval, knowledge tools, unverified freshness, gateway forwarding.
- Workflow schedules, durable waits.
- Tenant memory configuration, media.net configuration.
- Memory architecture and residual gates, evaluation floors, retrieval benchmark.
- Live issue states checked: Memory Loop PRD #2037 — closed, delivery tracker #2063 — closed, sampled audit #2654 — open, extraction billing #2669 — open, Team Knowledge #2995 — open, database-role cutover #3259 — open.