"I've had a lot of hits so you'd think I'd know in advance which ones will become hits. Songs that I was sure would become hits went nowhere and some songs that I didn't think anything of became my biggest hits"
It's been a long time since I heard this so I'm probably mangling it. A quick search shows that he probably did say something like this though https://www.birminghammail.co.uk/news/showbiz-tv/sir-elton-j...
The lesson is - just put it out there and see what happens
It's not something I necessarily associate with "good writing". Or does the skill transcend the medium and help you write better long form content as well?
Because if you can do that, you'll also be able to express yourself within character limits
Also - purely because you brought generations up: Twitter is a millennial product. the brainrot Gen is mostly Alpha, and some zoomers.
It's my personal experience at least, and since then I've never used social media attention as a gauge of the quality of my writing.
> Semoi is a plugin (currently only available for Obsidian) which tracks the length of time it took for a document to be written up
Trying to mechanistically prove that a human created some content as opposed to ai, in the age of LLMs and style transfer when you can just ask for something to be written in the style of Mark Twain or drawn in a style of van Gogh and get a great output, is a fool's errand.
All solutions to this end are going to be some form of attestation.
Even the proposed approach of tracking keystrokes and timing as a form of mechanical attestation, is going to be short lived because someone will train an AI on a corpus of human keystrokes and get a replication. May not even need an ai for this, a stochastic program could conceivably reproduce this behavior.
But as you say, we don't know how to write a checker for that...
Luckily, it is fundamentally easier to measure the quality of the writing and argument than it is to determine whether a human wrote a text or not in the face of motivated "attackers" with access to whichever judging mechanism you use.
The problem will be for the people who do care whether AI wrote something or not rather than about the quality.
You can remotely attest the input devices. You can do it anonymously (long story, but doable) and without requiring some kind of pre-signed image for the while OS. (The OS passes through recent-input attestations.)
For example, English is not my native language. I can speak, read, write - I have no issue using it for work or everyday life. My own kid only speaks English. But when we talk about writing an article, I would want to polish it. I would want to put my thoughts into a better form, so people may enjoy reading a well-written text which may have some fragments written or edited by AI so it will be simply better. I would use it for additional fact-check. Communication is not a competition in language skills.
I once heard a story told by a journalist. He used to write articles for the NYT from time to time, and the process was like this: he knows English, but the NYT asked him to write in his native language, a very experienced translator produced the English text, and then they polished it together with the editors. Once the article was published, it mentioned only his name — no mentions of editors or translators. Why? Because creating a text is not just writing or typing. Often it is a more complicated process which may or may not include other people or systems. What matters most is whether the author puts their signature at the end or not.
https://github.com/humthentic/itypedmypaper-v1
As others have pointed out, it's relatively a lot of effort to create an artifact that realistically current systems can pretty well forge.
I don't know that there is a scalable and comfortable solution to this problem (or at least one that is scalable and comfortable proportional to the demand for it).
I don't have any suggestions. I worry that the only strong solutions require a lot of power to be given to a centralized authority.
I'm not sure how you would actually do any of that attestation, if it's even possible. Text seems especially difficult. Maybe photo/video could be achieved with specialized hardware and cryptographic signing though I don't know much about either so I'm not sure how it would work. Maybe all that attestation would also tie in to some kind of universal personal identifier online, so that bad actors can be tracked or excluded and can't repeatedly spin up new accounts.
It might mean a big reduction in privacy for certain online spaces that opt in to such a system... but the alternative of all trust being eroded and voices drowned out by a sea of bots or generated content seems potentially worse.
Being published on the platform is the proof of authenticity, not a png that can be copied, or a link that the reader has to check.
"Authorship creates an originality report you can share with your professor, boss, or editor. It also lets you replay your document’s creation from the first word to the last punctuation mark so you can double-check your process."
#!/usr/bin/env bash
while true; do
git add -A
if ! git diff --cached --quiet; then
git commit -m "$(date --iso-8601=seconds)"
fi
sleep 5
doneAnd this doesn't even have to deal with all the stuff about mouse movements that you don't register?
Of course maybe I am just being typical programmer here, I guess lots of the people use generative AI would be defeated by copying pasting in the text and getting labeled AI, but that would also incorrectly label lots of people who have old texts in handwritten form they do not want to type all over again (of which I am one), and finally I assume that there is money in the field so producing something that allows bots to display "human heuristics" would probably get made and be profitable.
Funnily enough when I was automating things, generally twitter, I discovered that my real usage often got registered as bot, so I figured what's the use.
Also talking with someone who actually worked on bot-recognition by usage metrics said I was overly paranoid on some of the things I made my scripts do to appear human.
What matters a lot more is to stop abusive behaviour and low quality slop - whether or not that is written by humans or AI.
And working to detect that is a lot more tractable problem than trying to stop automation.
This means the certificate is independent of author and source text.
There's nothing stopping you from sending fake counts/duration to the semoi server. It's a certificate that only says "at this point in time, this is the information I was provided with".
You can then attach it to any piece of text you like.
At the very least, you'd need the ability to prove that there is an underlying event stream with these characteristics, and that this exact event stream creates the document in question. You still can fake that event stream, but it becomes enough work to distract at least casual abusers.
But really, it's the equivalent of saying "I wrote this without AI, honest" in-doc and signing that with your personal key. The value depends entirely on your willingess to be truthful. (IOW: I predict we'll see a resurgence of reputation systems, to some extent)
I think you're right that reputation systems are the best solution to a low signal-to-noise ratio.
Consumers (and the agents under their control) will increasingly prioritize content and other products with reliable attestation that they come from a trustworthy publisher, organization, brand, or individual.
We believe you on internet karma?
If all we're left with is "trust me bro", then we're left with nothing.
If someone writes "this was completely hand-written with no AI assistance", I'll just believe them. I'm already committed to letting your words fill my brain for a bit, so I don't know why I would NEED a cryptographic signature to PROVE you aren't lying about WHO wrote it.
Being called out as a lier will be a LOT more painful than 1) using an LLM to write quickly and not lying, or 2) doing what you claim to be doing and writing it with your feeble, non-metallic human hands.
(First version of this comment had an example from an HN thread of this SPECIFIC behaviour getting called out but ehh that’s not the right vibe. My point is it does happen.)
So the issue isn't "did a human write something", it's what the actual content is
The bonus, you can use whatever detection method as a judge for a coding assistant loop.
"/goal write a script to pass [some AI detector]"
Also, I noticed that taking a neutral stance in a discussion almost always got flagged as AI. To avoid being labeled AI, you had to have a strongly opinionated tone. But that naturally exposes weaknesses in the writing.
I think if the author is trustworthy, even AI-assisted editing can still be valuable.
These are the signs. Spotting AI writings isn't as hard as people think. AI is not the savior you think it is.
I think it's important to consider the risks/rewards/benefits. There's definitely a sense in my circle of contacts that if a work is seen as purely human then it's somehow better and more authentic, and annecdotally, it's also possible to win points by taking something created with the help of AI and passing it off as your own independent work. Like social credit, there's a sense that you'll seem smarter than you feel yourself to be.
With that in mind, absolutely any technical solution to detect AI or attest to human authorship will be abused. The only context that a solution like this one supports is one where there isn't a risk or reward, the author just sincerely wants you to know that it's human-produced.
Plenty of people, after all, produce a draft with AI, then type it all out again editing and refining and updating as they go, and then ask an AI to look at the result and make suggestions, and then go back and make the changes they agree with. Such a workflow would be deemed human by semoi - so it's lucky that there's no point in lying about it.
if it doesnt take off, it dies.
if it does take off and becomes a relevant currency of some sort, it will need improving. if it needs improving how far are we willing to go? does some alternative system fork off to handle severity of provenance concern?
what happens when AI action becomes indistinguishable from human action? what happens when sticking computers in your head becomes vogue?
if it takes off and fills a small niche, maybe thats the best future.
Like I said, cool idea, but seemingly very cursed from the get go.
There's plenty of microbehavioral analysis we can do that is initially effective but will get bypassed (with GANs being the purest way, or something more domain-specific).
You could imagine a livestream that's permanently published somewhere. But the verification of the livestream takes longer than reading the piece itself (and is itself vulnerable to faking).
Personally I think it comes back to something like writing under your real name - staking your reputation - as the most trustworthy indicator. At least people in your circles can trust you.
That means this basically works, as long as it never gets large enough to attack. Which may suit the author just fine. Not everything has to solve the world's problems. But it won't generalize very far, no.
I mean, I also think it's really pathetic to have an AI write something and then say "this text was written with no AI assistance"†. So if we've acknowledged that we're only going to stop non-pathetic people, why not skip the cryptographic hash signing and just go with the no-AI statement?
I understand that the goal is only to make lying hard, not impossible. However, I don't think this solution makes lying harder enough to meaningful.
I know it's inevitable, but I just hope we can collectively delay the end a bit longer.
I don’t think security through shame is practical.
Perhaps there might be some "metric" that can be calculated automatically from git commit history, but it would be nicer if the author can manually allocate "weight" to each passage of the text, based on what they perceive is most important (read the dark regions for the the deepest insights).
--
Use Case 1: Emphasis in office communications. When employee A communicates with employee B, they can use the text background to communicate the energy invested by employee A to produce the email/report/memo/doc in question.
Example 1: Developer spent 1 month rewriting the login system to a clean, feature-identical, drop-in replacement Auth API. When he announces to the CTO, the text is in solid grey, communicating the number of hours that went into this project and hence it's importance (better read all the details). dark highlight = I invested a lot of energy in this; I know you're a business person now, but I need you to read this and act on it!
(current alternative way to "highlight" actions is to link them to financial benefits for the company, e.g. the new Auth API will save us hours of dev time and allow us to integrate with platforms X, Y, Z, but I think the "number of hours I put into this" would be a good metric to show as well. It's proximal to the developer's emotional investment in the thing)
--
Use Case 2: Emphasize key ideas in educational text. Educators can communicate the relative importance of different parts of the text. Instead of "energy that went into creating this text," the more useful signal would be to tell readers how much mental energy it will take them to understand a given concept. A sentence that explains some key concept can be shown in dark background to say "this is deep" and also "this is worth learning about specifically within this section." Example 2: An author explains a sequence of 10 complicated steps that are part of some process, and indicates by text background color which steps are important and which are just technical minutiae.
* I write math and science textbooks, and I'd love to have such a sidechannel to the reader to communicate "importance" ... I hack by marking some key stences as bold, but it could be better.
--
Use Case 3: Ethical labelling of genAI output. It is now commonplace within the academic world to require disclosure of genAI use when creating scientific publications and other docs. This is a good practice, but it leaves too much degrees of freedom to the author about the level of disclosure they make. A word-for-word provenance metadata "channel" in parallel with the final text of the document could be a very useful thing to have (in an ideal world). In the real world, few academic would admit to heavily leaning on genAI, but at least we can shame them for not providing the "detailed provenance" track along with their text.
In contrast, people who are using genAI unashamedly (half the population) could prove how un-ashamed of their genAI usage they are by specifying "All genAI" in the "provenance channel" for all the posts they publish. If you like the SLOP or you think you can control the SLOP, I won't judge you, but please let me know so I read the text differently...
Example 3: Author A asks editor E to review a book draft. The manuscript clearly shows which parts of the book were written by the author and which parts are genAI. The editor knows which parts to focus their attention on (what the author is saying), and which parts are just filler. For bonus points, the author could also disclose the harness+context+prompt they used to produce the text (like a view source affordance that comes with any genAI text passage).
It's so difficult to write, you review, edit, edit and you read it and you can't unsee the LLM criticism.
ai detectors always have so many false positives that im convinced they're only put in place by people too dumb to know any better and snakeoil salesmen.
A few years ago I started writing Twitter threads [0]. A few weeks ago I passed 200 total threads.
When I started writing them, my thought process was: "Is anyone going to be interested in my stories/ideas??"
Dear HN comment reader, I can 100% assure you of two things:
1. If you write things, at least one person will read them.
2. It is VERY hard to predict what people will find interesting
e.g. some of the threads I thought people would find the least interesting got the most traction and vice versa. The only way to find out is to write it down.
I would also add that just writing, a LOT, helps you become a better thinker and writer. Twitter threads in particular are great as they force you to distill a story down into bite sized chunks.
One additional benefit: you meet amazing people when you write about what you are interested in. Why? Because if someone likes your writing, they would probably like talking to you and you to them.
0 - https://x.com/alexpotato/status/2012723178577985948?s=20