Tech
How to Build AI Agents With Memory Using Weaviate Engram
LLMs are stateless. Every API call starts cold. That works for one-shot answers, and fails for agents that must remember preferences, past decisions, and lessons across sessions.
Weaviate Engram is a managed memory service built on Weaviate for exactly that problem. You send raw conversations or events. Engram extracts structured memories, reconciles them with what it already knows, and stores them for semantic search. Your agent stays fast because memory work runs asynchronously, while recall stays precise because retrieval is backed by Weaviate’s vector index.
This guide shows how to wire Engram into a real agent loop.
Why agents need Engram (not just a bigger context window)
Stuffing full chat history into every request looks simple. It does not scale.
- Long context raises cost and latency on every turn.
- Models still get lost in the middle.
- Raw transcripts are noisy, contradictory, and outdated.
- Multi-agent workflows split one task across multiple windows, so “one transcript” is not enough.
Engram’s model is different: actively maintain memories. Extract facts. Deduplicate. Update when preferences change. Retrieve only what is relevant for the next decision.
What Engram is
Engram is a memory server for LLM agents and apps. It exposes a REST API (https://api.engram.weaviate.io) and a Python SDK (weaviate-engram).
Core capabilities:
Core concepts (keep these straight)
- Memories — discrete facts, embedded as vectors for search
- Topics — categories that guide extraction (e.g. UserKnowledge, experience)
- Groups — bundles of topics + a pipeline for one use case (often default)
- Scopes — who a memory belongs to:
- project-wide (shared learning)
- user-scoped (hard isolation via multi-tenancy)
- property-scoped (e.g. one summary per conversation_id)
- Pipelines — async graphs that extract, reconcile, and commit
Templates like Personalization get you started without designing pipelines from scratch.
Setup
- Create an Engram project in Weaviate Cloud (Personalization template is a good start).
- Create an API key and save it immediately.
- Install the client:
pip install weaviate-engram anthropic
# or: uv add weaviate-engram
export ENGRAM_API_KEY=”eng_…”
export ANTHROPIC_API_KEY=”sk-ant-…”
import os
from engram import EngramClient
client = EngramClient(api_key=os.environ[“ENGRAM_API_KEY”])
The agent memory loop
A practical agent loop with Engram has three steps each turn:
- Recall — search memories for the current user message
- Act — call the LLM with recent turns + recalled context
- Remember — fire-and-forget the new exchange into Engram
1) Store conversations (async)
run = client.memories.add(
[
{“role”: “user”, “content”: “I just moved to Berlin and prefer specialty coffee, not chains.”},
{“role”: “assistant”, “content”: “Got it — I’ll keep specialty spots in Berlin in mind.”},
],
user_id=”alice”,
group=”default”,
)
print(run.run_id, run.status)
Engram returns a run_id immediately. The pipeline:
- Extract — pull topic-matching facts
- Transform — dedupe / merge with existing memories
- Commit — persist to Weaviate
You can poll with client.runs.wait(run.run_id) when you need consistency before the next search. In most chat UIs, fire-and-forget is fine because the latest turn is already in short-term context.
Other input types:
- String — app events (“User viewed pricing page”)
- Pre-extracted — agent decides what to remember via tool calls
2) Recall before the model responds
from engram import HybridRetrieval
results = client.memories.search(
query=”What kind of coffee does the user like?”,
user_id=”alice”,
group=”default”,
retrieval_config=HybridRetrieval(limit=5),
)
memory_context = “\n”.join(f”- {m.content}” for m in results)
Retrieval options:
Minimal memory-enabled agent
import os
import anthropic
from engram import EngramClient, HybridRetrieval
engram = EngramClient(api_key=os.environ[“ENGRAM_API_KEY”])
llm = anthropic.Anthropic()
user_id = “alice”
recent = [] # short-term: last few turns only
def agent_turn(user_input: str) -> str:
# 1) Recall long-term memory
results = engram.memories.search(
query=user_input,
user_id=user_id,
group=”default”,
retrieval_config=HybridRetrieval(limit=5),
)
memory_context = “\n”.join(f”- {m.content}” for m in results) or “- (none yet)”
system = f”””You are a helpful agent with persistent memory.
What you remember about this user:
{memory_context}
Use memories when relevant. Do not invent facts not present here or in the chat.”””
recent.append({“role”: “user”, “content”: user_input})
# 2) Act with short-term context + recalled memory
response = llm.messages.create(
model=”claude-sonnet-4-5-20250929″,
max_tokens=1024,
system=system,
messages=recent[-6:], # last ~3 exchanges
)
assistant = response.content[0].text
recent.append({“role”: “assistant”, “content”: assistant})
# 3) Remember asynchronously
engram.memories.add(
[recent[-2], recent[-1]],
user_id=user_id,
group=”default”,
)
return assistant
This pattern replaces growing history with search + a small recent window, which cuts tokens while keeping personalization.
Give the agent control with tools
Automatic recall before every turn is simple. Tool-based recall is more powerful for multi-step agents.
Expose Engram as tools:
This matches the Hermes Agent plugin model (engram_search, engram_store, engram_fetch).
Sketch:
tools = [
{
“name”: “search_memory”,
“description”: “Search long-term memories about the current user.”,
“input_schema”: {
“type”: “object”,
“properties”: {“query”: {“type”: “string”}},
“required”: [“query”],
},
},
{
“name”: “store_memory”,
“description”: “Store or correct a fact about the user.”,
“input_schema”: {
“type”: “object”,
“properties”: {“content”: {“type”: “string”}},
“required”: [“content”],
},
},
]
def handle_tool(name: str, args: dict, user_id: str):
if name == “search_memory”:
return [
m.content
for m in engram.memories.search(
query=args[“query”],
user_id=user_id,
retrieval_config=HybridRetrieval(limit=5),
)
]
if name == “store_memory”:
run = engram.memories.add(args[“content”], user_id=user_id)
return {“run_id”: run.run_id, “status”: run.status}
When the agent “forgets,” it stores a correcting memory. Engram’s reconcile pipeline supersedes the old one instead of leaving contradictions in the store.
Continual learning for agents (not only users)
Engram is not limited to user preferences. Configure topics like experience or feedback so agents learn workflows over time:
- User says genre filtering should use a genres property, not near-text search.
- Engram extracts feedback, transforms it into an experience memory, and commits it.
- Next task, the agent searches experience memories and avoids the same mistake.
Scope choices matter:
- Project-wide experience — team agents improve together
- User-scoped experience — personal agents that never leak learning across users
Design patterns that work in production
- Always pass user_id for user-scoped topics — Engram enforces isolation; do not invent a shared memory bag.
- Use hybrid search by default — best balance of meaning and exact terms.
- Keep short-term history short — last 2–3 exchanges + recalled memories.
- Fire-and-forget adds; wait only when needed — e.g. before a critical next-step search.
- Use bounded topics for profiles — one UserProfile per user, fetched into the system prompt every turn.
- Let agents store corrections — do not delete as the primary “forget”; reconcile instead.
- Separate groups by use case — personalization vs continual learning stay clean.
REST fallback (any language)
curl -X POST “https://api.engram.weaviate.io/v1/memories” \
-H “Authorization: Bearer $ENGRAM_API_KEY” \
-H “Content-Type: application/json” \
-d ‘{
“input”: {“string”: {“content”: [“The user prefers dark mode.”]}},
“user_id”: “alice”
}’
curl -X POST “https://api.engram.weaviate.io/v1/memories/search” \
-H “Authorization: Bearer $ENGRAM_API_KEY” \
-H “Content-Type: application/json” \
-d ‘{
“query”: “What UI preferences does the user have?”,
“user_id”: “alice”,
“retrieval_config”: {“retrieval_type”: “hybrid”, “limit”: 5}
}’
Summary
Building agents with memory is not “save the transcript.” It is extract, reconcile, scope, and retrieve.
With Weaviate Engram you get:
- A low-latency write path (memories.add) that pipelines extraction in the background
- Weaviate-backed search (vector / bm25 / hybrid) for relevant recall
- Hard multi-tenant isolation by user and soft isolation by properties
- Two integration styles: auto-recall into the prompt, or agent-controlled tools
Start with the Personalization template, wire the search → respond → store loop, then add tool-based recall and experience topics as your agent grows.
Tech
iPhone 18 Pro series India sale begins today: Price, deals
The iPhone 18 Pro and Pro Max went on sale in India today, with confirmed pricing starting at Rs 1,64,900 for the Pro and Rs 1,79,900 for the Pro Max, available via Apple’s store, Flipkart, Amazon and retailers.
Apple’s iPhone 18 Pro series went on sale in India today, with the company confirming official prices for both the Pro and Pro Max models.
The iPhone 18 Pro starts at Rs 1,64,900 for the 256GB model, rising through Rs 1,89,900, Rs 2,39,900 and up to Rs 3,14,900 for the 2TB variant.
Sales began at 8 am IST via Apple’s official India store, with the devices also available on Flipkart, Amazon and other authorized retailers.
Those who pre-ordered ahead of the launch are being prioritized for delivery, while walk-in buyers can purchase from remaining retail stock.
Apple is offering cashback, EMI and trade-in options on the new phones, along with free Personal Setup for direct purchases.
The iPhone 18 Pro and Pro Max were originally unveiled globally at Apple’s September 9 event, alongside the AirPods 5, in the company’s first major launch under new CEO John Ternus.
The devices feature a 2nm A20 Pro chip, a 48MP Fusion camera system with variable aperture, and up to 30 hours of battery life on the Pro Max variant.
For the first time, Apple is offering a 2TB storage option on the iPhone 18 Pro lineup, aimed at users with heavy photo and video storage needs.
India’s official pricing was confirmed only closer to the local launch; earlier estimates ahead of the global unveiling had pegged the India price near Rs 1,39,900, lower than the confirmed figure.
Physical store availability is expected to remain limited in the initial days, with Apple prioritizing fulfillment of existing pre-orders before opening walk-in sales more broadly.
The iPhone 18 series does not include a foldable iPhone, despite months of speculation ahead of the global launch event.
India remains one of the fastest-growing markets for Apple globally, with the company expanding its retail footprint and local manufacturing base in the country in recent years.
Apple Store sales floor (representative image), Wikimedia Commons, CC BY-SA 4.0
Tech
Modi kicks off SEMICON India 2026 at Yashobhoomi Convention Centre
PM Modi inaugurated the fifth edition of SEMICON India 2026 at Yashobhoomi, New Delhi, with over 600 companies from 52 countries taking part in the three-day semiconductor industry event.
PM Narendra Modi kicked off SEMICON India 2026 at the Yashobhoomi Convention Centre in New Delhi, marking the fifth edition of the country’s flagship semiconductor industry event.
The three-day convention, themed ‘Silicon to Systems: Building the Ecosystem,’ brings together over 600 companies from 52 nations and more than 150 speakers across a 15,000 square metre exhibition.
The event’s opening day fell on PM Modi’s 76th birthday and Vishwakarma Jayanti, which he referenced in connecting India’s craftsmanship legacy with its semiconductor push.
A roundtable with global semiconductor CEOs ahead of the launch saw Japan’s Fujifilm commit ₹800 crore to a new semiconductor materials facility in India.
India’s Semicon India Programme has approved 12 projects to date, part of a broader push to build the country’s semiconductor market to $200 billion.
The government has approved 12 semiconductor projects under the Semicon India Programme so far, with three units under the initial Semicon 1.0 phase having already commenced commercial production.
Twenty-four startups have been approved under the Design Linked Incentive scheme, which supports Indian companies developing semiconductor chip designs.
The government has also approved a Semicon 2.0 phase of the programme, structured around six pillars aimed at deepening India’s semiconductor ecosystem.
Ahead of the inauguration, PM Modi chaired a roundtable with global semiconductor CEOs, emphasizing close collaboration between government and industry on skilling, innovation and emerging technologies such as AI and quantum computing.
Japanese company Fujifilm announced an investment of ₹800 crore for a new greenfield semiconductor materials manufacturing facility in India, following its CEO’s attendance at the PM’s roundtable.
The government has set a broader target of building India’s semiconductor market into a $200 billion industry as part of its long-term electronics manufacturing strategy.
September 17 also marked PM Modi’s 76th birthday, as well as Vishwakarma Jayanti, which the PM referenced in his address while linking India’s traditional craftsmanship to its modern technology ambitions.
Semiconductor wafer on dicing tape (representative image), Wikimedia Commons, CC BY-SA 4.0
Tech
Poco X8 Power, X8 5G launched in India, price starts Rs 29,999
Poco has launched the X8 Power 5G with a 10,000mAh battery at Rs 35,999, alongside the standard X8 5G at Rs 29,999, with sales beginning September 11 on Flipkart.
Poco X8 Power and Poco X8 5G have been launched in India, with prices starting at Rs 29,999.
The X8 Power, priced at Rs 35,999, is built around a 10,000mAh battery paired with 100W HyperCharge fast charging.
The X8 Power also features a Qualcomm Snapdragon 6 Gen 5 chipset and a 6.83-inch 1.5K, 120Hz AMOLED display.
The standard X8 5G runs on a Snapdragon 6s Gen 4 chipset, with a 9,000mAh battery and a 50-megapixel dual rear camera setup.
Both phones will be available from September 11, 2026, exclusively through Flipkart.
At 229 grams and 8.65mm thick, the X8 Power’s weight and dimensions reflect the trade-offs of housing a 10,000mAh battery in a phone of this size.
The standard Poco X8 5G runs on a Qualcomm Snapdragon 6s Gen 4 chipset, with RAM options of 6GB, 8GB and 12GB.
Camera hardware on the Poco X8 5G includes a 50-megapixel primary sensor paired with a 2-megapixel secondary lens, along with a 16-megapixel front camera.
The Poco X8 Power’s 10,000mAh Silicon-Carbon battery is paired with 100W HyperCharge fast charging, aiming to offset the downside of a larger battery with quicker top-up times.
Poco has positioned both phones squarely in the budget-to-midrange segment, competing against similarly specced offerings from Redmi, Realme and other Chinese-origin brands in India.
Both phones will go on sale from September 11, 2026, at 12 noon, exclusively through Flipkart.
Launch offers include a Rs 2,000 coupon discount and a Rs 1,000 bank offer or exchange bonus, bringing effective prices down to Rs 26,999 for the X8 and Rs 32,999 for the X8 Power.
The Poco X8 Power 5G is powered by a Qualcomm Snapdragon 6 Gen 5 chipset and features a 6.83-inch 1.5K, 120Hz AMOLED display.
The X8 Power carries IP66, IP68, IP69 and IP69K ratings for dust and water resistance, an unusually comprehensive set of protections for its price segment.
Wikimedia Commons, CC BY-SA 4.0 (representative Poco smartphone image)
-
Brandpost2 years agoRedfox Overseas: Customize Your Own Energy Drink and Stand Out in the Market
-
Brandpost2 years agoZamzam Company CEO Chhote Bhai-Bade Bhai gave a grand welcome to Indian writer Devhari Sirvi in Dubai
-
Fashion2 years agoShikha Sharma: The Fashion Journalist, Blogger, and Plus-Size Model Taking the Industry by Storm
-
Entertainment2 years ago
Lucky Roxx’s “Yadav Ki Pukar” Goes Viral! Youth Celebrate Unity & Power
-
Entertainment9 years agoNew Season 8 Walking Dead trailer flashes forward in time
-
Entertainment9 years agoMeet Superman’s grandfather in new trailer for Krypton
-
Uncategorized2 years ago
Hello world!
-
Brandpost2 years agoFrom Local Hero to National Icon: Satyam Yadav’s Journey with TwentyOne
