The memory layer
your AI is missing.

TEMPRA is an agnostic, on-premise memory infrastructure that plugs into any LLM stack — no re-architecting, no third-party cloud, your data fully sovereign.

AI with a native sense of time.
Not memory bolted on — time encoded as geometry inside the embedding space.


Powered by the Temporal Embedding Memory Projection Reconsolidation Architecture
Below: a glimpse of the future consumer app — on the roadmap. The infrastructure ships first.
9:41•••
TEMPRA · TEV

Your threads

H Thesis Alexandre Cannae Elephants Rome Kant
247 memories · 431 links
9:41•••
Wednesday · April 20

Hello Sofia

RESUMING · YESTERDAY, 10:14 PM
"Compare Hannibal to Alexander, on military strategy…"
2 points remaining: Cannae 216 av. J-C, Alpine elephants.
SUGGESTIONS
PROJECTHistory project pitch
STUDYReading · Kant
9:41•••
Close ● LISTENING · 00:42
"…yes, I remember — Mathieu prefers red,so a light pinot will be perfect."
Infrastructure-first · B2B

One memory layer. Any AI stack.

TEMPRA is not a SaaS and not a chatbot. It is agnostic memory infrastructure you deploy inside your own environment — and plug into whatever you already run.

Model-agnostic

Works behind any LLM — Claude, OpenAI, Mistral, Llama, your fine-tunes. No vendor lock-in.

On-premise & sovereign

Runs in your VPC or on your own hardware. Your data never leaves your walls.

Drop-in API & SDK

A single integration point. Plug it into your apps, agents and pipelines in days — not quarters.

Native temporal geometry

Time (τ) is encoded inside the embedding space via the TEV — not bolted on as metadata.

Governance & forgetting

Selective, cascading forgetting and full audit control over exactly what is retained.

Built for scale

Vector-native retrieval (HNSW / pgvector) engineered for production loads and large memory graphs.

How it works

Three steps. One memory that lasts.

Where bolt-on memory forgets, TEMPRA gives any AI stack a continuous temporal thread — model-agnostic, deployed in your own environment.
01 — CAPTURE

It listens without forcing.

Your apps and agents stream conversations, documents and events. TEMPRA extracts the facts that matter — not raw logs.

02 — WEAVE

TEMPRA connects the threads.

The engine builds a living temporal graph of entities, events and projects across your data — noting what repeats, contradicts and refines.

03 — RECALL

It surfaces at the right moment.

Query months later through one API. TEMPRA returns context with its sources — the right memory, at the right time.

March Today
TEMPRA · The engine

Memory built to last.

T.E.M.P.R.A (Temporal Embedding Memory Projection Reconsolidation Architecture) is the system that makes TEMPRA a long-term companion. It doesn't store everything — it retains what matters, with time as a native coordinate.

  • Semantic extraction: it keeps facts, not raw transcripts.
  • Relational graph: each memory points to the people, places, and projects involved.
  • Contextual recall: memory surfaces only when it's useful.
  • Selective forgetting: you decide what stays, what goes, what gets pinned.
  • Sovereign by design: memories live inside your infrastructure, never ours.
TEMPRA habits people projects places
Capabilities

Everything a memory layer should be.

Everything your team needs to ship AI with real memory — and nothing that locks you in.

Time-aware recall

Retrieve the right memory for the right moment with time-aware ranking — not flat vector similarity.

Memory graph

A queryable graph of entities, events and projects. Inspect links, pin, retrieve, export — full transparency.

Data control

Zero training on your data. Encryption in transit and at rest. Export or purge anything, on demand.

Ingest anything

PDFs, transcripts, tickets, emails. TEMPRA absorbs them and remembers across months — still in context.

Drift & contradiction

TEMPRA flags when state changes over time — what an entity or user believed then versus now.

Scoped agent memory

Give each agent its own scoped memory, all drawing from one shared, governed temporal graph.

Why TEMPRA

Your AI already works.
But its memory is bolted on.

RAG and vector stores retrofit memory from the outside. Here's what changes when time lives inside the model's memory.
Bolt-on memory — RAG & vector stores
Memory sits outside the model

Context is fetched and stuffed into the prompt on every call. The model never truly holds it.

Flat facts, no time

"User lives in Paris." No timeline, no decay, no evolution. A snapshot, frozen forever.

You maintain the plumbing

Chunking, embeddings, re-rankers, glue code — your team owns the whole retrieval stack.

Stale state goes unnoticed

When facts change, retrieval still surfaces the old ones. No notion of what is current.

With TEMPRA
Memory lives in the embedding space

Time (τ) is a native coordinate via the TEV — not metadata fetched at query time.

A timeline, not a notepad

Every memory is timestamped and weighted. March-state and today-state are different — and TEMPRA knows it.

One layer, inside your stack

A drop-in API, deployed on-prem. No chunking pipelines to babysit, no glue code to maintain.

It tracks drift and contradiction

"In November this was X — now it's Y. Here's what shifted." State that evolves, surfaced automatically.

Bolt-on memory retrieves.
TEMPRA remembers — time as native geometry, not metadata stapled to a vector store.

The principle

Most systems bolt memory onto the model and call it remembering. TEMPRA makes time part of the geometry — so the memory is the model's, not a database it queries.

TEMPRA — The Living State
Carthage AI
FAQ

The most asked questions.

How is this different from RAG or a vector database?

RAG fetches text and stuffs it into the prompt at query time — memory stays outside the model. TEMPRA encodes time as a native coordinate inside the embedding space (the TEV), so recall is temporal and intrinsic, not a similarity lookup bolted on afterwards.

Can we control what is retained?

Yes. Every memory can be pinned, edited, or forgotten. There's also a "selective forgetting" mode that lets TEMPRA unlearn an entire theme in cascade — linked people, projects, associated conversations.

How do we deploy it?

On-premise, in your own VPC, or as a managed deployment — your choice. TEMPRA is infrastructure: it runs where your data already lives, with no third-party cloud in the path.

Is there a consumer app?

A consumer app is on the roadmap. Carthage AI is infrastructure-first: the agnostic B2B memory layer ships before the app, which will later be built on top of the very same engine.