Faster prompt lookup drafting in llama.cpp

https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp/

Comments

jadidbourbakiSep 27, 2026, 3:29 AM
Btw, if anyone has experience with the open source community in general and llama.cpp in specific, I would greatly appreciate some advice. I’m facing a bit of an interpersonal issue that I really hope is resolved without any ill will. Here is the context:

https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/comment...

Any advice for what I can do? Due to this, I cannot create a PR or issue in the llama.cpp repository. However, I am worried about bothering the maintainers on other channels in case it aggravates them further. Thank you for your help!

electroglyphSep 27, 2026, 7:07 PM
hopefully someone on the team sees this and fixes the situation, otherwise, if it's good, i'm sure somebody else will open the PR. thanks for your contribution.
wronglebowskiSep 27, 2026, 8:39 PM
I understand this repository must be under siege effectively with how popular it is, but the team needs to come up with a decent way to handle this. The current climate with Nvidia buying Hugging Face who effectively owns the llama.cpp team there’s a lot of discourse around their control of the team that’s negative.
embedding-shapeSep 27, 2026, 9:21 PM
> Any advice for what I can do?

Email that author (email can usually be found via GitHub, or contact via other private way they've shared somewhere) and explain the situation, don't lambast them publicly on social media or similar ways. If that doesn't work, do the same but for another maintainer. Don't spam all of them straight up, wait a week or something before contacting someone else.

netburstSep 28, 2026, 12:56 AM
That would be inadvisable as the maintainer (?) in question's GitHub bio states:

> Emailing me about a temporary ban will result in a permanent one instead.

As far as I can tell the ban was given based on a single accidental ping to the maintainer, who then reviewed the draft PR before realizing it was not for the mainline llama.cpp repo.

rfgplkSep 27, 2026, 8:37 PM
Just hard fork the project. Frankly, llama.cpp is so badly written that these kind of speedups are trivial, and a hard fork (or a total rewrite) has been needed for the longest time.
jadidbourbakiSep 27, 2026, 12:20 AM
Fun update to this: Daniel Lemire added another optimization to make this even faster. https://github.com/jadidbourbaki/llama.cpp/pull/12

I’ll benchmark his change and add it to the article, crediting him for this improvement.

jadidbourbakiSep 28, 2026, 3:13 AM
As an update to this, Lemire absolutely crushed it! I benchmarked his change today. Here is my PR comment:

https://github.com/jadidbourbaki/llama.cpp/pull/12#issuecomm...

S0ySep 27, 2026, 11:39 PM
How much does this speedup inference for the end user in terms of tk/s ?
PicardsFluteSep 27, 2026, 7:41 PM
So, instead of resolving it privately like you were asked, you come here and complain? Was the bot posts on Reddit not enough for you? Pro tip: There is a right way and a wrong way to approach these things and you are MOST certainly approaching this the wrong way. And before you go accusing me of anything: I am NOT associated with that project, but I DID see what happened. You are acting like a damned child and should be ashamed of yourself.
ricardobeatSep 27, 2026, 8:32 PM
Is there some missing context here? Where was the author asked to resolve issues in private? And what 'private' channels would be available to them?
frazar0Sep 27, 2026, 9:16 PM
I think they meant to reply to this comment

https://news.ycombinator.com/item?id=49859982#49863097

but commented on the post.