
Who Athena-Public is for#
Developers maintaining long-running codebases across AI sessions
Anyone working through hundreds of AI sessions on the same codebase benefits most from the /start and /end lifecycle. Past decisions, architecture choices, and debugging sessions are available in the next session without pasting context manually.
Skip if:
If you are working on a greenfield project or a short task contained in a single session, the /start and /end discipline adds overhead without proportional benefit.
Developers who switch LLMs or IDEs frequently
Because Athena's state is plain Markdown on disk, switching from Claude to Gemini to a local model preserves all accumulated session context. There is no per-provider memory setup; the same files work with every supported model.
Skip if:
If you use one model in one IDE and have no plans to switch, the portability benefit is smaller, though the session compression and hybrid retrieval still apply.
Developers evaluating AI agent memory frameworks
Athena publishes its benchmark methodology and 644-test suite openly and labels each feature as code-enforced, agent-discretion, or aspirational. For engineers who want an accurate picture of what a memory layer delivers before adopting it, the repo's transparency is notable in the category.
Skip if:
If you need a memory layer that works with no configuration, Athena requires a one-time Python setup that some native IDE memory features skip entirely.
Solo developers and small teams building internal AI tooling
The zero-cloud-cost baseline (local compute plus Supabase free tier, S$0.00/month) and MIT license make Athena viable for personal or internal projects without budget overhead. The starter kit includes 426 pre-built protocols covering common engineering workflows.
Skip if:
Teams that need centralized memory shared across multiple developers will need Supabase for cloud sync; the default local path is per-machine and not automatically shared.
The problem it solves#
AI coding agents have no long-term memory. Each session starts from zero: the model has no idea you decided on a specific architecture in session 14, debugged a race condition in session 47, or established conventions the team agreed on in session 83. You re-explain the same project constraints, solve the same problems twice, and watch months of accumulated context disappear with every model update or session reset.
The situation worsens when you switch models or IDEs. Platform memory systems like ChatGPT Memory, Claude Projects, and Gemini Gems store flat facts in vendor-controlled storage you cannot edit, export, or move. When a model update resets a personality or you switch to a different provider, you start from zero again. Developers working on codebases spanning hundreds of sessions have no durable way to carry what the agent knows from one session to the next.
How it solves it#
Session lifecycle with /start and /end
Typing `/start` in your IDE's chat panel loads roughly 2,000 tokens of compressed project context into the agent's window. `/end` saves the session back to your Markdown files. Three modes are available: Lightweight (~500 tokens for quick questions), Standard (~2K-10K for daily use), and Deep (~20K tokens for complex planning).
Hybrid RAG retrieval
Queries use BM25 keyword matching combined with pgvector semantic search and a cross-encoder reranker via Reciprocal Rank Fusion. On a 65-query gold set, retrieval Hit@5 is 0.569 and MRR@5 is 0.472. The README publishes these strict figures and explains why an earlier lenient metric of 0.892 was deprecated, noting it counted partial substring matches as hits.
Seven-IDE portability
Works with Claude Code, Antigravity, Cursor, Gemini CLI, VS Code with Copilot, Kilo Code, and Roo Code. Each IDE gets a config file (CLAUDE.md, AGENTS.md, .cursor/rules.md, and others); the underlying Markdown state is identical across all of them. Switching IDEs means pointing a different client at the same files.
Path-triggered skill injection
Context is loaded based on which files are active in the current session, not by dumping the full project state into every query. This targeted injection saves 80-98% of context window tokens compared to loading everything upfront, reducing per-session model costs.
Governed autonomy with code-enforced hooks
Hooks block destructive commands before execution in code, not as a prompt instruction to the agent. The ruin check blocks 14 of 16 destructive command patterns; two known bypasses are tracked in the repo's tech debt log. Secret scanning runs in CI on every commit, with zero leaks across 1,248 commits.
Zero-cost local baseline
The baseline configuration runs entirely on local compute plus a Supabase free tier. Infrastructure cost is S$0.00 per month. A full install adds cloud sync and cross-encoder reranking. The PyPI package (`pip install athena-agent`) installs the lightweight version with no cloud dependencies.
Strengths and trade-offs#
Strengths
- Plain Markdown state you own and can version-controlContext lives as plain Markdown files on your disk that you can read, edit, diff, and git-track. When a model is deprecated or you switch providers, you point a different client at the same files. Nothing is locked inside a vendor's opaque storage.
- Model and IDE agnostic by designAthena works with Claude, Gemini, GPT, and any LLM accessible from the supported IDEs. The README describes the design as 'rent the intelligence, own the state.' Switching LLM providers does not require re-establishing context; the files work the same way across all models.
- Published benchmarks with honest methodologyThe README reports strict Hit@5 (0.569) rather than a lenient metric (0.892) that counted partial substring matches as hits. The repo also documents known bypasses in the ruin check and labels each feature as code-enforced, agent-discretion, or aspirational, a level of transparency rare in the AI agent category.
- 644 passing tests and zero secret leaksThe test suite runs 644 unit and integration tests at 100% pass rate. Gitleaks scans every commit, with zero secrets detected across 1,248 commits. CodeQL and a privacy gate also run in CI on every push.
Trade-offs
- -Compounding personalization is single-author evidenceThe README flags the 'session 500 recalls session 5' claim as N=1 evidence; 1,900+ sessions by a single author with no multi-user study. Developers building Athena into team-scale memory workflows are running ahead of the documented evidence base.
- -Anti-sycophancy gate is Claude Code onlyThe meta-awareness gate that prevents the agent from silently increasing agreement over time is code-enforced only in the Claude Code integration. In all other supported IDEs, this relies on agent discretion rather than a hard code check.
- -Setup requires Python, a virtual environment, and IDE configurationGetting started requires cloning the repo, creating a Python virtual environment, running `pip install`, and running `athena init` for your specific IDE. The README's quickstart guides this in about five minutes, but it is not a zero-configuration install.
Athena-Public vs alternatives#
Athena vs Pieces for Developers
Both tools target AI-assisted developer workflows, but they address different problems. Pieces for Developers focuses on capturing and retrieving code snippets and context fragments from your IDE in real time; Athena focuses on maintaining compressed project state across AI sessions through a structured loop.
| Feature | Athena | Pieces for Developers |
|---|---|---|
| License | MIT | Proprietary |
| Self-hosting | Yes, local compute | No, cloud-managed |
| Storage format | Plain Markdown on disk | Proprietary cloud |
| Memory model | Session state via /start and /end | Snippet capture and retrieval |
| LLM support | Any (Claude, Gemini, GPT, local models) | Integrated with specific providers |
| Cloud dependency | Optional (Supabase free tier) | Required |
| Cost | Free forever (self-hosted) | Freemium with paid tiers |
Athena is the better fit when you want long-term project context that survives model switches, full data ownership with no vendor dependency, or the ability to inspect and edit your memory files directly. The session lifecycle builds up compressed context across hundreds of sessions, and the hybrid retrieval (BM25 plus semantic search) surfaces past decisions without manual search. The MIT license means you can modify the tool and use it commercially with no restrictions.
Pieces for Developers is worth considering when you need frictionless, passive snippet capture from your IDE without any setup routine or session discipline. If you want context captured automatically rather than structured deliberately, Pieces handles that without requiring the /start and /end conventions Athena depends on. Pieces also has a polished UI and supports more platforms, including mobile, for accessing snippets across devices.
Quick start#
Deploy Athena by cloning the repo and running the Python install and setup commands for your IDE.
```bash
git clone https://github.com/winstonkoh87/Athena-Public.git && cd Athena-Public
python3 -m venv .venv && source .venv/bin/activate
pip install -e ".[local]"
athena init --ide claude
athena doctor
```What it's built on#
- Languages
- Python
FAQ#
Does Athena work with models other than Claude?
Yes. Athena is model-agnostic and works with Claude, Gemini, GPT, and any LLM accessible from the supported IDEs. Context is stored as plain Markdown files on your disk, so you point a different model at the same files. Tested IDEs include Claude Code, Antigravity, Cursor, Gemini CLI, VS Code with Copilot, Kilo Code, and Roo Code.
How is Athena different from native memory in ChatGPT or Claude Projects?
Native memory features store flat facts in vendor-controlled storage you cannot edit, export, or move between providers. Athena stores session context as plain Markdown files on your disk. You own the files, can inspect or edit them, and can point them at a different model when a provider updates its memory policies or you want to switch tools.
What is the retrieval accuracy of Athena's hybrid search?
On a 65-query gold set evaluated against the author's session history, Hit@5 is 0.569 (strict) and MRR@5 is 0.472. The README documents the methodology and why an earlier lenient metric of 0.892 was deprecated; it had counted partial substring matches as hits, inflating the score without improving real retrieval quality.
Does Athena send data to external services?
The baseline configuration sends no data to cloud storage. All processing runs on local compute. An optional Supabase integration is available for cloud sync of vector memories; if you use it, data is stored in your own Supabase project, not in Athena's infrastructure.
Is Athena ready for team use?
The README labels the compounding personalization evidence as N=1, based on 1,900+ sessions by a single author with no multi-user study. The anti-sycophancy gate is code-enforced only in the Claude Code integration. Athena works well for individual developers today; teams using it for shared memory should treat it as a tool with known single-author evidence and test it before relying on it.
Similar open-source tools#
rakazo
AI teammates you own: your keys, your model, your machine.
Hermes Agent
Self-hosted AI agent with persistent memory, multi-channel chat, and model choice across OpenAI, OpenRouter, and custom endpoints.
Litellm
Self-hosted AI gateway for 100+ LLMs in OpenAI format
Pstack Claude
Structured coding workflows for Claude Code, Codex, and Pi
Recall
Local project memory for Claude Code, generated entirely offline.
Gstack
Turn Claude Code into a virtual engineering team with 23 skills

