I showed them an innocuous working example of how they could classify thousands of participant entries as "mentioning Italian food" whether it mentioned "pasta", "penne", "rigatone", etc etc. Jev did this task with high confidence, and showed it could differentiate between the inverse, eg had very low confidence when I switch entry to mentioning "hamburgers". I find that a very useful tool to pass on to non-technical colleagues.
But on HN, there are lots of much more qualified people saying Jev is a con, a step backwards, or bonkers that people are impressed. Will one of them tell me plainly what software was doing this prior to Jev, and the like? I don't doubt it existed, but I never came across it. I'm interested to hear from people more expert than myself.
To be precise, software must be able to: - Classify thousands of entries with arbitrary themes or topics, even if entry doesn't mention that theme or topic directly, with high enough confidence to be useful. - Do it with thousands of entries in 0-3seconds. - Do it at negligible cost, eg sub 0.1cents
I was inspired by self-extracting archives. I wanted to share files with basically no dependencies. The goal was:
1. Something that didn't require any installation (assuming a web browser)
2. Have a single file with no network that could self-decrypt
3. Be fully auditable
The second point is done by having (sort of(*)) reproducible builds and embedded OpenPGP signatures.The first point is made by cleverly manipulating the HTML structure so that it can decrypt without breaking the PGP signature. It can even decrypt using bare openssl (which was a design goal too, though getting the exact structure right took some work and bug reports).
The third point is accomplished by the first two, and by the source being freely available.
(*) Depends on the OS at the moment.
I just open sourced the DSL that our harness in grep.ai uses to turn repeatable parts of agent work into workflows. You can combine tool calls, code, Jev-powered system one decisions for things like routing and screening evidence, and agents when a step needs more investigation.
Our harness uses the traces and retro notes agents leave behind when doing a job to figure out which parts can become a workflow. The idea is to make the work easier to understand and avoid paying for a full agent loop where one isn’t needed. For example, a research workflow can split a question into subquestions, send agents to research them in parallel, use Jev to screen the evidence, and have another agent write the report. You can inspect the steps, evaluate the evidence screening separately, or change one agent without rebuilding everything.
The DSL and examples are in our GitHub. There’s a scripted demo you can run without API keys: https://github.com/Parcha-ai/agentrun
You can also use it as a Pi extension to build, inspect, and run workflows: https://github.com/Parcha-ai/agentrun#use-it-in-pi
I would love to hear if this is useful to others.
More background on how AgentRun works in this video: https://www.youtube.com/watch?v=vOVhtGjtwpg. Or read about our use cases in this article: https://x.com/MiguelriosEN/status/2101029313906987422.
This last part, along with everything else this past year, has caused me a great deal of stress.
While our balcony zen garden project is yet to be completed, I had an idea to create a virtual one that everyone can use.
It's an idea that, unfortunately, I’ve been postponing for a while now, mostly because I have no fucking idea how to do it as I don’t know JavaScript, and I don't have the capacity to learn it right now.
So I let perfect be the enemy of good and, well... just kept the idea to myself.
Then I said "fuck it" and used AI to make the thing I really wanted to make.
I realized I didn't want "perfect". I wanted "good enough".
I tweaked, added, removed, drew, researched, questioned, tested... I just wasn't the one coding it.
So now, instead of occupying my brain, it now lives on the internet for others to enjoy.
Yes, there’s something noble about making something entirely on your own, but what good is an idea that just sits in my head?
So here I am. I made the thing. The weird, little, quiet koi pond.
The silly project of passion. The little corner of the internet to let strangers watch fish quietly, together.
I hope this pond helps you as much as it helped me.
Find duplicate code blocks meaningful to the language (classes/functions), not just lines.
Find near-duplicates: ignore whitespace, strings, high level AST nodes such as function and names.
Find structurally similar code: anonymize identifiers, constants, etc.
Pull requests welcome: This is very much an proof of concept - I'm happy with it, but I haven't supported very many languages at present.
Languages supported: astro, bash, css, go, html, javascript, lua, markdown (plus codeblocks), python, sql, typescript, java, kotlin, rust, yaml