Posted in Content.
Can you prove you writing is your own, and it wasn’t AI-generated?
Written by
Daniel Prindii
on .
Miscellanea is the newsletter for the independent thinker, at the intersection of digital content, information management and culture. In this edition, I discuss how to have a history of your work and the costs that GenAI puts on knowledge workers.
Edition No 15. Date: 25 August 2026
If you like this newsletter, please forward it to a friend. You can sign up here.
Last month, a publishing deal for a crime novel was cancelled after the literary agents lost faith in the author’s use of GenAI in the writing and editing of the novel. Earlier this year, the winning short story of the prize from Caribbean on Saturday faced accusations of being written with GenAI. Last year, Deloitte was under fire for submitting two government reports citing AI-hallucinated research.
Keeping track of the authorship of one’s content is becoming increasingly important, both from a perspective of proving this is your work (not AI-generated slop), and from a critical thinking one (is the information fact-checked and valid). So what is happening in this space?
Anthropic’s Claude is watermarking the content generated by its models. The system will not include hidden characters, but it will change “the source of the randomness used to pick among words”. As LLMs generate responses word by word, “it chooses among a list of potential candidates, ultimately selecting the most sensible or likely based on the preceding text”. What I’m reading between the lines is that AI will sound more like AI, but in a watermarked, proprietary way.
Of course, accuracy is the “hot” topic here. As with the recent Substack partnership with Pangram, the same piece of writing can be “AI-generated” or human-generated.
Independent research (linked in Recommendations) that tested the difference between LLM-generated and human-generated words found:
“Findings show that while detection tools can provide useful initial flags, they should not be used as sole evidence in high-stakes decision-making but should be implemented in a broader evaluation strategy.”
Another study (linked in Recommendations) showed that ChatGPT is biased against non-native English speakers: “While the detectors accurately classified the US student essays, they incorrectly labelled more than half of the TOEFL essays as "AI-generated" (average false-positive rate: 61.3%).”
Plus, every LLM provider has its own way of creating watermarking and I don’t expect them to share the details so that ChatGPT can identify Claude’s output.
These changes will be implemented by all providers, because they signed the EU AI Act, where the Article 50(2) Code of Practice on Transparency of AI-Generated Content describes commitments, how to make them work and limitations.
The Article introduces transparency in four situations:
when people interact directly with AI
when AI generates content
when AI is used for biometric categorisation or similar recognition
and when deepfakes are created with AI on matters of public interest.
From a platform perspective, it’s hard to see how AI-generated content is going to be fairly labelled. The law is casting a wide net, we don’t have a general agreement on what counts as synthetic content or how to measure it, and our AI literacy, when it comes to recognising it, is unreliable. Offloading the effort to users, like LinkedIn is doing with its recent “Looks like AI” button, is just a witch hunt waiting to happen.
In a YouGov survey conducted on behalf of Pangram, 35% of all American adults couldn’t differentiate between human and synthetic content.
And from a people perspective, declaring the AI use relies on users’ good faith. I bet you already thought of at least one use case where someone had all the reasons to use AI and not admit they did it.
So how can you document that your writing is human-generated, so you can prove the quality of your work? The solution I see is to create an archive of your research and keep a journal of the progress you make.
The checklist
These points will give you enough structure right from the start and cover the most possible scenarios.
1. Write with digital tools that have a version history and can be easily synced and backed up.
2. Save your drafts: have a folder for digital drafts and one for your paper ones.
3. Archive and document everything you do: voice memos, screenshots, paper notes, discussions. Make notes in the research journal, add a date, a tag, and see what connection you can make with the rest of the research. Allocate energy and time for doing this, even if you feel it’s boring. New ideas can emerge while doing paperwork.
4. Disclose and explain the use of GenAI in the project, if applicable.
5. Cite your sources.
The tools
There are tools to help you have an organised archive of your work.
The Obsidian note-taking app. It’s local-first, so everything is saved on your computer. You can sync your files and create backups with third-party tools. You can check the system I have in Obsidian in this article. It covers the technical setup and how I’m organising the information inside.
Ellipsus, a cloud-based writing editor. Created and hosted in Germany, it has a clear no-AI policy. On their paid plan, they have Emboss, an authorship tool to share writing metrics, versioned history, sessions, and other metrics recorded in the doc.
Your standard office suites. LibreOffice, the free and open source office suite, has a version history for its documents. The same is true for Google Docs, Microsoft Office, and Proton Docs.
Of course, there is always the mighty pen and paper combo: private by default and resistent in time. Except for cats, water and fire. If you go this route, consider digitising with some scans or photos of the work done.
The emotional cost and value of the process
The annoying part is that the burden of proof for documenting the process falls on you, the writer. For a small piece, like a blog post, the work around the work is not much. A few notes, an outline, some research, quotes and links. But what happens when you scale? What about a thesis? Or a research project with more collaborators and authors? How do you keep all in check and organised?
The whole creative process will receive a similar status as the sketches and work processes of artists. Version histories will show train of thought, provenance, process and creative connections of the final work. Will we be having a market (market as capital and market as audience) for these process details? Or will we be keeping them only so we can prove to an all-knowing Bob from the platforms that we wrote those books? My bet is that the process is going to become just as valuable as the end product.
Currently reading
I’ve started Dune by Frank Herbert. Here’s a very contemporaneous quote: «“Once men turned their thinking over to machines in the hope that this would set them free. But that only permitted other men with machines to enslave them. Thou shalt not make a machine in the likeness of a man’s mind,” Paul quoted.»
Recommendations
The research cited before: “How are combinations of human-written words and LLM-generated words by ChatGPT, Copilot, Gemini and Grammarly detected by Turnitin?” link
The research, published in 2023, about the bias in generative models: “# GPT detectors are biased against non-native English writers” link
The German folk-doom band Empyrium and its 1997 album Songs of Moors and Misty Fields. Perfect to get you in an autumnal mood. Spotify
Main image by Arie Oldman on Unsplash
Tagged
You might also like
-
How to run an independent newsletter
- Written by Daniel Prindii
- Published on
-
On Liquid Content and LLMs
- Written by Daniel Prindii
- Published on
-
Summer reads recommendations
- Written by Daniel Prindii
- Published on