Second Brain
GBrain needs to be tweaked
The best second brain is the one that is scalable and fits your needs.
Marco Rodrigues · September 1, 2026
Large Language Models (LLMs) are becoming more intelligent. Agent harnesses like Claude, Codex, and Cursor are becoming more robust, with many different features. Overall, AI is becoming easier to use, with almost no friction to get started and increasing productivity.
However, no major competitor in the field has come up with a reliable second-brain option. People are still using Obsidian, LLM Wiki from Andrej Karpathy, and GBrain from Garry Tan.
There are other solutions out there, but in the end, they all rely on the same technology: either a folder of Markdown files, a vector database, or both.
Take Obsidian and GBrain, for instance. They are both called second brains, but in fact, their purposes are quite different. The first acts more like a notebook, while the second is a retrieval engine.
Many people try to oversell Obsidian, but in the end, it is a text editor with links for Markdown files sitting in a folder. Its primary features are visualization, portability, and searchability through frontmatter. But it is not made for agents.
On the other end, GBrain keeps the Markdown files as the source of truth, but adds a real engine. On every page written, three things happen:
- The page is embedded. Text is chunked and turned into vector embeddings, then stored in the vector database. In addition, it uses HNSW (Hierarchical Navigable Small World) for fast similarity search.
- The page is keyword-indexed. The full text goes into a column, so BM25 keyword matching works on exact names, code identifiers, and phrases.
- The page is linked into a graph. References to entities such as people, projects, and tools are extracted and connected through typed relationships, allowing the system to understand how different pieces of knowledge relate to one another.
Therefore, for anyone looking for a more agentic second brain, GBrain is better suited because it can act faster, more accurately, and spend fewer tokens.
However, there's a problem with GBrain. It was not made for everyone. It was made for Garry Tan, the creator, and that poses an adoption issue.
This piece explores the architecture of GBrain and the kinds of tweaks that can make it work better for different needs, along with the pros and cons of changing the core GitHub repository.
Markdowns and backlinks are not enough
Obsidian is a collection of Markdown files that contain links to others.
For example, imagine there is a note about machine learning and another note about chemistry. In the machine-learning note, [[Chemistry]] can be written. Obsidian treats this as a link to the Chemistry note. More specific notes can also be linked, such as [[Neural Networks]] or [[Drug Discovery]]. Over time, these links create a network between the notes.
And the backlink is the other side of that link. If the Machine Learning note contains [[Chemistry]], then the Chemistry note can show that Machine Learning links to it. As more notes link to Chemistry, Obsidian collects those incoming links and shows them as backlinks.
This makes it possible to open the Chemistry note and immediately see which other notes are connected to it. This is great from a visualization point of view, but not as useful for context and information retrieval.
In brief, this works for humans, but not for agents. And as the brain grows, not even humans.
Adding a vector database makes this much more useful because it allows agents to search for concepts rather than relying only on explicit links or exact keywords.
Meet GBrain and its architecture
GBrain is an open-source knowledge-graph engine for AI agents built by Garry Tan, Y Combinator's CEO. It was originally made for himself but was eventually released for others to use.
Its GitHub repository already has tens of thousands of stars and thousands of forks, showing the level of interest around the project.
In brief, the system is based on a directory of Markdown files, similar to Obsidian, but the files get parsed, embedded, linked, and queried through a Postgres (PGLite) database.
GBrain has its own CLI, so it can also be used without an AI agent. These are the key commands that illustrate how the system works:
- gbrain sync — scans the Markdown repository and imports new or changed pages into the database. It parses the frontmatter and page content, tracks changes incrementally, and, for smaller changesets, can generate embeddings as part of the import process. In other words, this is what keeps the database in sync with the Markdown files.
- gbrain embed — generates vector embeddings for the chunks of text that don't have them yet. The Markdown content is split into smaller chunks, converted into embeddings, and stored for semantic search. This is particularly important after a large sync or when pages were imported without embeddings.
- gbrain extract — builds additional structure from the Markdown. It can extract links between entities and create timeline entries from dated information. These relationships are then stored in GBrain's knowledge graph, allowing the system to understand connections between people, companies, projects, events, and other entities without needing an LLM call.
- gbrain dream — runs the background "thinking" or enrichment cycle. Instead of simply storing what is already in the Markdown files, it periodically looks across the brain to synthesize and improve the information: consolidating knowledge, finding patterns, fixing or enriching pages, and surfacing things such as contradictions or missing information. This is what makes GBrain feel more like a continuously evolving memory rather than a static database.
- gbrain search — queries the brain using hybrid search rather than relying on a single retrieval method. It combines vector similarity with keyword search (BM25), then applies additional ranking and reranking techniques to find the most relevant pages.
The system has other nuances that are not mentioned in the list above, but those commands are the core of a functional GBrain. Now, it is worth taking a look at the architecture:
- Vector Database: The default is PGLite, which runs in-process, so there is no server or Docker setup required, and it is recommended for smaller brains, roughly up to 50K pages. For larger setups or multi-machine sync, it can use regular Postgres with pgvector, either through Supabase or a self-hosted instance. Multiple agents can query the same brain.
- Embeddings: GBrain uses embeddings to represent the meaning of the content and make semantic search possible. The default for new installations is Voyage voyage-4 with 1024-dimensional embeddings, although GBrain supports several different embedding providers. For reranking, it defaults to Voyage rerank-2.5.
- Hybrid search: Search does not depend on embeddings alone. GBrain combines vector search using HNSW, traditional BM25 keyword search, and reciprocal-rank fusion to combine the results. It then applies a cross-encoder reranker to improve the final ranking. There are three search modes: conservative, balanced, and tokenmax.
- Knowledge graph and schemas: GBrain adds a type system on top of the graph, so entities can be classified as people, companies, projects, events, and other domain-specific types. There are several schema packs that can be selected.
- Autopilot: This is a built-in background loop that can continuously work on the brain without requiring an external scheduler. It also includes a durable job queue for running subagents and shell jobs, allowing some of the enrichment and maintenance work to happen in the background.
- CLI and MCP interface: GBrain has its own CLI for interacting with the brain directly, while also exposing most of the same operations through MCP for AI agents. Some operations remain CLI-only, particularly those that interact directly with the local filesystem.
Taken together, these components make GBrain more robust than other so-called second brains. Because it can easily scale without compromising retrieval speed and token costs, it can become particularly useful when an agent is involved.
With that being said, should GBrain simply be installed and left alone? It is not so easy, because it may not be tailored to every use case.
The original GBrain was forked, and you should too
Now that GBrain's architecture is understood, it becomes easier to explain why changes to the original repository can sometimes be necessary.
Database: From PGLite to PostgreSQL
The default GBrain setup uses PGLite, an embedded version of Postgres compiled to WebAssembly. It is a great default because it requires no server, no Docker, and almost no setup. However, it is not as scalable.
That's one of the reasons Garry Tan mentions Supabase for a more scalable approach. Another option is to start with a faster and local approach: PostgreSQL.
These are the two main reasons that can make such a switch useful:
- Performance and Speed: As the brain grows and graph queries become more expensive, PGLite's WebAssembly-based execution can become a bottleneck.
- Concurrency: GBrain uses multiple processes running at the same time: live sync, email and calendar imports, dream cycles, health checks, and link extraction. But PGLite can only write one process to the database at a time.
The tradeoff is that this change breaks the Autopilot feature of GBrain. That means the automation needs to be rebuilt separately.
Schema: From generic to custom
The schema is another important part of GBrain. It tells the system what kinds of things exist in the brain and how they can be connected.
The default gbrain-base schema defines 22 page types, including:
- person
- company
- deal
- project
- meeting
- slack
- calendar-event
- concept
- media
- source
- writing
- analysis
- code
- image
- diary
- event
- note
It also defines typed relationships such as works_at, invested_in, founded, attended, and led_round.
This is important because it explains the relationships between pages and what they are. The schema can also use frontmatter and language patterns to infer some of these relationships automatically. If a person page has company: SpaceX, for example, GBrain can turn that into a works_at relationship in the graph.
The 22 types in Garry Tan's schema are generally enough, but a shorter and more tailored list can make sense for specific needs. A custom schema can instead use types such as finance, hiring, task, idea, research, personal, and household, along with additional typed relationships.
This change can create conflicts with how the system is built, but those issues can be solved afterward. Such changes may initially appear lighter than they actually are, because modifying the schema can affect other parts of the system.
Model Strategy: Different Models for different jobs
The agent harness used with GBrain can be the Hermes Agent by the Nous Research team. While Garry Tan appears to favor Claude, an open-source solution can provide a different approach.
Using one model for everything is also not necessarily efficient.
For instance, the dream cycle can run across thousands of pages, so using an expensive reasoning model for every operation can quickly become expensive. A better strategy is to use cheaper models for high-volume tasks and reserve stronger models for work that actually requires more intelligence.
For example:
- Embeddings: OpenAI text-embedding-3-large
- Bulk extraction and synthesis: DeepSeek V4 Pro 0813
- Reasoning-heavy tasks: Grok 4.6
- Heavy enrichment: GPT-5.6 Sol
In the case of the Hermes Agent, auxiliary models can also be used to save even more tokens in the process.
Automation: From autopilot to cron jobs
As mentioned earlier, changing the vector database to PostgreSQL can break the Autopilot mode, which runs synchronization, dreaming, and maintenance in the background.
Instead, the Hermes Agent can be used to run the brain through explicit cron jobs:
- Live Markdown: database sync every 15 minutes
- Gmail and Calendar imports: every 6 hours
- An overnight dream cycle
- Link extraction
- Health checks and automatic remediation
- Type-drift detection
- Meeting-note processing
- Person and company enrichment
- Graph quality checks
- Upstream update checks
This method has some benefits because it gives more control over almost every step of the brain's cycle. It becomes possible to decide exactly when something runs, which model it uses, and how much it is allowed to spend.
Ingestion: From a simple folder to automation
GBrain is Markdown-first when it comes to data ingestion. Therefore, gbrain sync keeps the database synchronized with the Markdown folder.
GBrain also provides ingestion recipes for external sources such as Google Gmail, Calendar and Contacts, meeting transcripts, X, and voice, but these integrations need to be configured separately.
For example, the Google integration can pull Gmail threads, Calendar events, and contacts. Calendar events can be converted into Markdown pages, which are then imported into GBrain and embedded so they become searchable.
Instead of configuring individual integrations whenever they are needed, a more automated brain can continuously ingest information from sources such as:
- Google Workspace — Gmail and Calendar
- Granola — meeting notes and participants
- Voice notes
- External enrichment — web, LinkedIn, X, and other sources
These pipelines convert the incoming information into Markdown and write it into ~/brain/. From there, the normal GBrain pipeline takes over: sync the files, embed the content, extract relationships, and make everything searchable.
Enrichment: Create Ideal Customer Profiles (ICP)
Instead of having a person page that contains a name, company, and a few notes, an enrichment pipeline can build a much richer profile covering things like:
- What they are building
- Their motivations and beliefs
- Events and presence
- Achievements and deals
- Network
- Trajectory
The information can come from multiple sources, including X, LinkedIn, web research, calendar history, and the existing brain.
This update turns a contact stub into a useful profile, making the graph much more valuable for actual decision-making.
Conclusion
This piece is not intended to criticize GBrain or any of the other so-called second brains. The main point is to show that:
- Markdown-only second brains like Obsidian are not made for scale or for agents.
- Hybrid second-brain approaches are better for information retrieval.
- GBrain may need some configuration to work for specific needs.
With that being said, GBrain can simply be cloned and connected to a preferred harness, but that by itself may not be enough.
That is why significant changes can be useful:
- Create a custom schema instead of using the generic one.
- Use native PostgreSQL instead of PGLite, mainly because it can be significantly faster and handle concurrent processes better.
- Automate ingestion from Gmail, Calendar, meetings, voice notes, and external sources, so the brain grows without relying on someone to manually upload the information.
- Use Hermes as the operator, handling ingestion, maintenance, health checks, and remediation.
- Add a custom enrichment layer that turns basic people and company pages into much richer profiles.
A customized version is not necessarily better than the original one. It is simply a fork made for a specific set of needs.
And that is how people should start looking at second brains.
Similar to how users pick their preferred harness — Claude, Cursor, Codex, or the Hermes Agent — or build a specialized agent, second brains are no different.
The existing options do not need to be treated as fixed. A solution can be taken and adapted to fit the requirements of a specific workflow.