Show HN: Minigraf – An embedded, bi-temporal graph database in Rust

github.com

29 points by adityamukho 10 hours ago

A successor project to my previous attempt at building a temporal graph database (https://news.ycombinator.com/item?id=23455516). This one builds on its predecessor's concept, but drops the dependency on another database to provide the storage, transaction and query engines. Minigraf is self-contained, packaged as a library crate, embeds into your application, and supports full bi-temporality.

real_faxenoff an hour ago

As I understand it, this is a tool for small datasets with an unknown schema and a required history? I'm struggling to find examples of real-world scenarios where this would be used.

LLM memory.. let's say, architecturally, its suitable for storing and processing queries - is questionable, but okay. But what real-world use case would generate such queries? It seems to me that the context for formulating and processing the results would be orders of magnitude greater than simply writing down what's needed in plain text. If my notebook took the form of a temporal multiverse of entities, I'd lose the ability to use it properly.

But maybe I'm just too old to quickly get the hang of every new type of Pokémon.

P.S. I'm using the built-in SQLite to store a graph with up to 10 million edges in an MCP app I'm developing for working with code.

eduardokairalla an hour ago

interesting project, but i don't understand when i should use minigraf over something like Neo4j, Cozo, or SQLite. can you clarify?

  • lovlar 8 minutes ago

    its a graphdb like neo4j and cozo and direct-to-file (instead of client-server with network delay inbetween) like sqlite.

    Also, I could be wrong but this project sounds like it has an index optimized for bi-temporal data (system time and real world time). in neo4j you would have to add that information as indexed properties on each node and edge in addition to other indexes (?).

    looks like this is an attempt to create efficient agent memory.

adsharma 7 hours ago

Why do you need another database and another query language for this?

Existing cypher based databases which support typed properties (including timestamps) on relationship tables can handle this use case just fine.

  • baq 5 hours ago

    bitemporal does not mean 'including timestamps' and out of existing RDBMSes only the truly most expensive ones support some kind of temporal queries; I've checked out for a bit, but if there's any with native bitemporal support that'd be quite the news

    • adsharma 4 hours ago

      Bi-temporal means the DB keeps track of transaction time and valid/effective time separately.

      Most databases have a notion of logical clock to implement MVCC/transactions. So this is just a mapping. Some in the spanner family actually use a drift limited physical clock.

      The second part is using an index to answer queries about facts in the past. Like what was the capital of India in 1800?

      There are existing embedded graph databases which do this.

      Disclosure: I maintain one.

    • Tanjreeve 4 hours ago

      Is bitemporal more than being able to time travel by transaction and having some sort of time lookup for next transaction? It's a new term to me.

  • Tanjreeve 6 hours ago

    Most graph databases I'm aware of require a server and aren't embedded into applications. Using datalog and applying it to graphs also more general and thus preferable than cypher which is still a niche specifically for graph databases.

    Also "because they can" seems obligatory.

  • ex1fm3ta 6 hours ago

    I feel you, I am completely lost with all new ''graph flavoured DB". I keep challenging those (with the help of Opus model ) and the answer is always the same "use progresql".

    • phoghed 5 hours ago

      Have to watch out that Claude hasn’t just figured out that you like being told to use Postgres lol

      • ex1fm3ta 4 hours ago

        Nope. Memory and data sharing for training are turned off. (I put 'turned off' in quotes because I don't know if Anthropic really keeps its promise. Often, it is cheaper for companies to break the law and just pay the fine.) I also use adversarial review and the '10th man' rule a lot. In fact, I spend more time planning than executing. I explicitly told Claude never to make a choice without proving it with numbers. By numbers, I mean benchmark results and score-based decisions (for example: 0.5 points for speed, 1 point for libraries, 2 points for documentation). The only thing I doubt are benchmark numbers, because they change depending on the hardware and what you test. Anyway, we always end up choosing PostgreSQL for large projects or SQLite for light ones (which makes sense).

    • Tanjreeve 4 hours ago

      Opus recommending a database server to swap in for an embedded file based database is either an outcome of the overwhelming volume of cloud based web dev + API glue code chat it was trained on or it's being obsequious but wrong.

fusslo 5 hours ago

what does 'Embedded' mean in this context?

forgive my ignorance on such a basic concept. I read 'Embedded' as in 'for Embedded systems'

Does it mean 'Embedded in your rust application'?

  • itishappy 4 hours ago

    Yes, "embedded database" is a term-of-art for in-application database. Many (most?) databases run a client-server architecture even for local apps.

  • andrashorvathde 4 hours ago

    We can think about a web application that can use SQLite or Postgres. Both use SQL as language, but Postgres is a separate service with its own port and ops work.

canadiantim 4 hours ago

How does it compare with the since-discontinued embedded graph database cozodb? https://github.com/cozodb/cozo

It also used datalog, written in rust

Great to see tho, I'm always eagerly hoping for a successful embedded graph database. Most recently my hope was for kuzudb, but they got acquihired; though it's still open-sourcing' as LadybugDB.

Highly recommend taking inspiration from Kuzudb/Ladybugdb and cozodb.

Godspeed!

  • adsharma 4 hours ago

    Maintainer of LadybugDB here. I'm reviewing a PR to transpile GQL to Cypher.

    It should be possible to transpile datalog or another query language that is interesting.

    Implementing storage and secondary indicies is the hard part.

    Won't comment on rewriting the DB in Rust beyond what's in GitHub discussions.

    • vkozio an hour ago

      Hi adsharma! Great to see you here! I've been thinking about a data query IR that databases could expose. There are IRs for algebra like Substrait, so maybe something similar could work here, maybe with binding info. I don't mean a low-level physical/plan IR but more like a "logical IR" that the optimizer can optimize :)

aavisangle 4 hours ago

what different features it offer than sqlight?