Our driver use case was development of specialist agents with a request/response lifecycle. Think of agents that have a well established input/output contract where they are expected to produce high-quality output (artifacts, responses etc.). Especially when the problem is in some verifiable domain and the LLM can iteratively refine a result to a final value that satisfies constraints or optimizes some goal.
The product is a coding agent skill + a serverless execution runtime with a harness that takes in a system prompt + Python functions as tools. The coding agent takes in the requirements from the user, and tries agent variants by executing prompt/tool variants it creates.
It works best for cases where you can think of how you can evaluate a candidate agent - when you describe this information to your coding agent, it often does a decent job at building candidate prompts, tools and even benchmarks.
Prompts and Python tools that the coding agent creates integrate with a harness that implements an FSM that is tuned to drive an iterative refinement process for verifiable domains. This tuning enables one to use small models like Luna to produce high quality results while keeping costs at a manageable level.
Our website is not 100% complete yet (some examples are missing write-ups, not all use cases we tried are there etc.), but the system is operational and docs are there.
Thanks!
It runs the two in series instead:
1. A pure-Python Rete engine evaluates YAML rules against your facts. The verdict comes only from here. Same facts, same verdict, every time, with salience-based conflict resolution. 2. RAG retrieves passages from your own policy documents, and an LLM writes a plain-English explanation of the decision that was already made, citing those passages. It can't change the verdict.
A few things that went further than I expected: - Rules are a graph, not flat lists: nested all/any/not, and rules can assert facts that other rules consume (forward chaining). The decision trace shows the causal chain. - Audit mode records every rule evaluated, including the ones that didn't fire, condition by condition, with a snapshot of the rule set for replay. - Rules can steer retrieval (a fired rule narrows which documents get searched), and retrieved text can be turned into facts for the engine. - Non-technical authors can build rules in a visual editor, or paste a policy document and get LLM-drafted rules with citations. Drafts are never saved without review. YAML is still there for engineers.
The landing page has a live demo with no signup (8 demo domains: loan, fraud, clinical, insurance, legal, ops, e-commerce, blockchain). There's also an MCP server, so Claude and other agents can call /decide as a tool: `uvx ai-rete-rag-mcp`.
To be upfront: it's a hosted product with a free tier. The MCP client is open source (MIT, github.com/zaharajabeen13-create/ai-rete-rag-mcp); the engine and platform are not open source right now.
I'd especially like to hear from anyone who has had to explain an automated decision to a regulator or an auditor: what did they actually ask for?
However, I quickly pivoted like a true founder to create an app that addressed a personal issue of mine.
On Android, if you have a bunch of apps and fitness monitors from different manufacturers, very few of them all talk to each other. Your data is scattered across different apps or double counted. However, most data sources write to Health Connect.
Hamilton is a dashboard of your Health Connect data with a homepage that shows your key metrics at a glance. Your choice of data source is flexible to avoid double counting, and you can overlay and offset metrics on charts to view correlations and track progress over time.
I named the app after my dog, who occasionally fetches things and is responsible for my daily step count.
Below is an FAQ, and to avoid confusion I have helpfully labelled my responses [APP] or [DOG]: https://news.ycombinator.com/item?id=49848948.
This was a pattern across more than a year, spanning multiple devices (all iphones). It happens on both wifi and cellular. It happens with wifi calling enabled and disabled. It happens with people I call every day, in my address book, as well as new contacts (say someone I met who called to follow up). I called support, searched online, reset network settings multiple times at different points months apart.
After being on T-Mobile for about a year, the degraded service is about the same as Verizon. Every week, more than once, people call me, my phone never rings, and I never see a missed call on my device. It happens at home, at work, on work trips in different states.
I still get missed calls regularly and my device shows I missed the call. It seems like a certain % of callers get my VM, but I never know the call attempt happened. The missed calls never appear later on my phone (like a late VM or text deliver, coming minutes or sometimes hours after it was sent).
Hypothesis: A growing number of cellular subscribers don't care enough or use voice calls as much, so (regardless of motive/feelings) cellular providers have found they can simply reduce quality without losing customers.
I've spent years debugging Windows crashes with tools that were either friendly but limited (e.g. Visual Studio) or powerful but archaic (e.g. WinDbg). I developed patterns and methods for understanding what was going on, and decided to build it into a much more effective debugging tool called ForensicDbg.
I built a modern interface to minimize the friction when debugging. All of the data shown to you is analyzed, interpreted, and presented to you clearly, so you can focus on what matters. Everything is interlinked so you can quickly and intuitivly navigate through the process space.
ForensicDbg comes with an MCP server which allows for agenic debugging. The work done to interpret and interlink your data also benefits AI tools. It removes the risk of hallucinations while building a stable foundation for them to work from without spending tokens.
If you want to try it out you can sign up and get a free beta license here: https://www.forensicdbg.com/beta