TEMPRA is an agnostic, on-premise memory infrastructure that plugs into any LLM stack — no re-architecting, no third-party cloud, your data fully sovereign.
AI with a native sense of time.
Not memory bolted on — time encoded as geometry inside the embedding space.
Works behind any LLM — Claude, OpenAI, Mistral, Llama, your fine-tunes. No vendor lock-in.
Runs in your VPC or on your own hardware. Your data never leaves your walls.
A single integration point. Plug it into your apps, agents and pipelines in days — not quarters.
Time (τ) is encoded inside the embedding space via the TEV — not bolted on as metadata.
Selective, cascading forgetting and full audit control over exactly what is retained.
Vector-native retrieval (HNSW / pgvector) engineered for production loads and large memory graphs.
Your apps and agents stream conversations, documents and events. TEMPRA extracts the facts that matter — not raw logs.
The engine builds a living temporal graph of entities, events and projects across your data — noting what repeats, contradicts and refines.
Query months later through one API. TEMPRA returns context with its sources — the right memory, at the right time.
T.E.M.P.R.A (Temporal Embedding Memory Projection Reconsolidation Architecture) is the system that makes TEMPRA a long-term companion. It doesn't store everything — it retains what matters, with time as a native coordinate.
Retrieve the right memory for the right moment with time-aware ranking — not flat vector similarity.
A queryable graph of entities, events and projects. Inspect links, pin, retrieve, export — full transparency.
Zero training on your data. Encryption in transit and at rest. Export or purge anything, on demand.
PDFs, transcripts, tickets, emails. TEMPRA absorbs them and remembers across months — still in context.
TEMPRA flags when state changes over time — what an entity or user believed then versus now.
Give each agent its own scoped memory, all drawing from one shared, governed temporal graph.
Context is fetched and stuffed into the prompt on every call. The model never truly holds it.
"User lives in Paris." No timeline, no decay, no evolution. A snapshot, frozen forever.
Chunking, embeddings, re-rankers, glue code — your team owns the whole retrieval stack.
When facts change, retrieval still surfaces the old ones. No notion of what is current.
Time (τ) is a native coordinate via the TEV — not metadata fetched at query time.
Every memory is timestamped and weighted. March-state and today-state are different — and TEMPRA knows it.
A drop-in API, deployed on-prem. No chunking pipelines to babysit, no glue code to maintain.
"In November this was X — now it's Y. Here's what shifted." State that evolves, surfaced automatically.
Bolt-on memory retrieves.
TEMPRA remembers — time as native geometry, not metadata stapled to a vector store.
Most systems bolt memory onto the model and call it remembering. TEMPRA makes time part of the geometry — so the memory is the model's, not a database it queries.
RAG fetches text and stuffs it into the prompt at query time — memory stays outside the model. TEMPRA encodes time as a native coordinate inside the embedding space (the TEV), so recall is temporal and intrinsic, not a similarity lookup bolted on afterwards.
Yes. Every memory can be pinned, edited, or forgotten. There's also a "selective forgetting" mode that lets TEMPRA unlearn an entire theme in cascade — linked people, projects, associated conversations.
On-premise, in your own VPC, or as a managed deployment — your choice. TEMPRA is infrastructure: it runs where your data already lives, with no third-party cloud in the path.
A consumer app is on the roadmap. Carthage AI is infrastructure-first: the agnostic B2B memory layer ships before the app, which will later be built on top of the very same engine.