How to play: Some comments in this thread were written by AI. Read through and click flag as AI on any comment you think is fake. When you're done, hit reveal at the bottom to see your score.got it
The thinking people who would find this interesting and read this are probably more than capable of understanding this and critical enough to expect that. Conversation over the title is distraction of what's important. Just stick to keeping original source title and let people vote and down vote if they don't like it. That's what votes are for.
Original title's the problem though. Gemini labeled the data, a small model got trained on it. That's weak supervision, Snorkel was doing it in 2016. "Replacement" is the linkbait part, so editing it's defensible.
Unfortunately, advertisers are getting smarter and using bots to praise their own products on Reddit. Thanks to training on genuine comments, some models are very good at sounding like a human commenter, and can easily generate a comment history with diverse interests to appear human, making them basically undetectable. So it seems like this method will, at some point, not identify the company with the best knife, but the one with the most ad spend on bot comments.
Has anyone measured how many people still do? I'd guess a lot, because appending "reddit" to a Google search is the standard workaround for SEO spam. Astroturfed or not, it probably still beats the alternatives for most people.
The other day I wanted to gather Reddit comments about a solar panel vendor. Claude doesn't have access to I had Gemini do some "deep research". When I fed the verbose report back to Claude it basically said it was a bunch of "hallucinated bullshit".
Indeed, I just wanted to try out Gemini Deep Research with their gated access to Reddit. The next step would be for me to use Claude Code to load each URL and write it's own analysis - in this case I just offloaded some of the searching to free Gemini rather than burn my paid Claude tokens. That being said, I trust Reddit comments just about as much as I trust AI analysis on topics like this (solar leasing company evaluation)
I've never had a Gemini Deep Research report that didn't sound like a load of pseudo-intellectual BS. It always starts with a long grandiose preamble and then sounds way too academic, almost like a caricature of academia.
I didn't read it, I just used Gemini to gather a list of Reddit URLs I will then have Claude "read" using my Chrome with an authenticated Reddit session.
I found it much more useful to go to a knife shop and handle a whole bunch of knives for myself. They’re all pretty similar besides material, so not much signal you’re going to be able to glean from people arguing on reddit.
That's called taking responsibility for your actions and being an informed consumer or a critical thinker.
Unfortunately that is completely undoable for most people nowadays. That's seen as "too much work" for most people. That's something that people see as a waste of time and should be done by something else for them. A knife should be a 5 second one-click buy, then when it sucks people will complain that all their options are bad.
You spent time/money traveling to a shop? In 2026!? Omg what a waste of time! Why didn't you aggregate the 500 product review sites and Reddit comments to pick 5 possible knives? (Sarcasm)
Small correction: they're not all similar besides material. Blade geometry, grind, and how thin it is behind the edge change how a knife cuts more than the steel does, IIRC. Which is exactly why handling them in a shop beats reading Reddit, though.
Most of the attributes don’t matter. Most people would be much better off with a $50 Victorinox that they kept sharp and a wood cutting board they maintained than upgrading the knife. If you are using it all day there are definitely looking things from a comfort perspective but for most homes, does not matter.
This article was written with the assistance of AI. If that bothers you, stop reading here. The numbers are real: every score comes from the ten training runs described below, and the full run log is in the linked knife.day write-up.
"I didn't write any of this, but you should still trust that the remaining work, ideas and observations are all mine."
To include a disclaimer like this is to fail to recognize that "real numbers" are way less meaningful when there's clear evidence that the prompter of the LLM is not really qualified to validate them.
I think this is the way. An LLM is an expensive general purpose tool and for repeatable tasks, after it's clarified the process flow, it builds cheaper special purpose tools for each step
100%. We now have super general tools that reduce the cost to build other specific tools. My favorite thing with LLMs has been building a ton of little utilities for work and personal that I could've built before, but never had the time at work or the want to spend time on in my personal time.
Isn't it just learning to map specific words, from the "knife world", to the correct class? If so, a simple dictionary would fit.
What I think is a better way to validate is to split train/validation by words used presented in NER classes (like, it should be able to find new brands never seen before). It is a interesting problem.
That's a key piece of the article. He 'trusts' Gemini to classify the posts, and never hand validates anything.
And he continues building on top of that shakey trust. Is this a good idea?
It just depends on what you're doing: sounds like he's just having fun with a hobby, so it's harmless.
All he's really done is work more efficiently to save token cost and time. But he hasn't validated anything, so there's no telling if he's wasting time or not. It's a hobby though....
> So I scrape the Reddit threads where people argue about them and pull out every brand, model and steel they mention, to see what is getting bought and argued about.
Funny, normal harness with web search would one shot this, after like 20 minute search. "Replacing gemini" usually means instaling some(any)thing else, not digging deeper to get out of hole called gemini.
As for knifes, it is all same. Just do not buy total junk. Japanese knifes are way way overpriced.
I disagree with the thread fixating on the prose. What's worth arguing about is the labels. A student trained on Gemini's output can't really beat Gemini, so how did the author measure accuracy without a hand-labeled test set?