Give Google the boot by building your own search engine
Got $10 and a gig of storage space? Marlin lets you weight your crawl toward the parts of the web you actually care about
software
Give Google the boot by building your own search engine
Got $10 and a gig of storage space? Marlin lets you weight your crawl toward the parts of the web you actually care about
If you're fed up with search results drowning out the bits of the web you actually care about, you could always build your own search index, as one developer did.
Nottingham, UK-based software dev Alex Morley-Finch built the open-source project dubbed Marlin for himself, cataloging around 560,000 homepages for roughly $10 in rented cloud GPU time and using less than a gigabyte of disk storage.
Morley-Finch said in a writeup of the project that he wanted a search engine of things he cared about, like “portfolios, zines, weird little art projects, one-person software,” and other cases of just “people doing stuff” that they share on the internet.
Something simple, with “a crawler that only ever looks at homepages, a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags,” he explained. “No IP scanning, no Redis, no storing full page HTML, no recrawl scheduler, nothing multi-tenant.”
Morley-Finch built Marlin around four processes: A fetcher that grabs domains, a worker that makes calls to a small, OpenAI-compatible language model, a steward that prevents bad pages from making their way into the index, and an API with a web UI where he can track the process and actually conduct searches of his index, complete with filters.
Of course, you can’t expect something like this to go perfectly on the first try.
“The first version worked within a couple hours. Point it at a sample of domains, watch things get summarised, search for them. Great,” he said in his writeup of the project. “Sunday afternoon [he began working on Marlin on a Sunday], I looked at what had actually been catalogued and it was the wrong web.”
Instead of indexing what he wanted it to, Marlin had just been grabbing a cross section of the internet, leading to more than 90 percent of the first attempt being “corporate sites and documentation.”
Rather than blocking certain domains, Morley-Finch built a weighting system to push certain pages to the top of his crawl queue, and others as far down the list as possible.
Morley-Finch rented a cloud GPU to handle most of the heavy processing and stopped the crawl at around 560,000 pages after exhausting his prioritized categories and beginning to pull in "the raw internet." He spent around $10 on the cloud GPU.
The one overarching problem he had, and which he said may be a problem for anyone else who tries to replicate his project, is tagging and categorizing - left to a language model, those important elements got a bit messy.
“I let the model invent its own category and tag names freely, on the theory that it would teach me the taxonomy instead of me guessing one upfront,” he explained, noting that it helped things get started quickly, but “I ended up with 671 distinct categories, many used exactly once, and over 121,000 tags, more than half used only a single time.”
Morley-Finch had to build a manual merge tool to clean up the mess, leading him to declare that his design likely won’t scale well beyond a hobbyist project, as “‘just let the model freestyle’ catches up with you fast.”
Still, he reckons Marlin’s rough edges shouldn’t put off anyone willing to spend a weekend tinkering with it.
“I think this is very doable in a weekend, and cheap enough that the GPU bill isn’t the reason not to try,” Morley-Finch wrote. “The code is going up as open source, so you can point your own crawl wherever your own curiosity leads.”
The Marlin GitHub repo is filled with how-tos and details for those who want to create their own weighting list based on their priorities, with the option to rent a cloud GPU if their hardware isn't up to the challenge. ®
Originally published on The Register
