I feel like there are probably some ad based search engines which are privacy and service oriented, but in general even for those there remains a misalignment problem. Hence if I don’t want to be a product now or in the future, what good search engines are there that I can pay for?


When I said ‘direct expenses’ I mostly meant the cost of owning / running a database of internet pages and metadata comprehensive enough to be considered part of a ‘fully featured search engine’. There’s also the other half; the compute required to create that metadata, as well as obtain it, but at most I would guess that those would be equal in cost to just having the space for a database of all the internet pages (scaling up after that based on how many users you need to support). In short, a scaled down web engine that had access to every page on the internet that people would want to find could cost as low as 100,000$ for a first time purchase for the hardware.
The internet archive does in fact have their own web crawler they use. They also do sites upon request as well; i’ve had my personal website on there for almost two decades now, specifically at my request.
They also have a full-featured search function available for anyone on their website at archive.org. This is why I say they’re a reasonable price comparison for a full-featured search engine. They may spend more on storage and less on metadata compute than a theoretical smaller search engine, but at the end of the day, that’s just a re-balancing of the cost, not a completely new and more excessive cost.
I think direct expenses; the cost of owning and maintaining an internet index database, are definitely significant enough that the completely free access that google gives to anyone who wants it, are way more than any single private entity or company is able to support just because they want to have it. I don’t think it would be anywhere even close to a billion dollars though.
I think the hardest part of having a internet index database would be the knowledge required to create and maintain it, especially under the hostile forces that are the 75 billion dollar seo industry. If a selfhosted search engine became big enough that the seo industry started trying to break it, I don’t think that company would survive for very long at all.
Google is losing that battle, like, almost completely. What hope would a small startup style company have of battling it and staying financially solvent, especially if they’re trying to be different from google and bing and actually showing results without the pressure of advertisers breathing down their necks?
I think the hardware side of a search engine is solvable with silicon valley startup level of funding. I think it’s impossible for anyone in the current day and age to make that sort of project solvent while keeping the user (instead of the advertiser) as the main customer. For anyone else who can’t get those funds, or don’t actually want to do a results-oriented search engine, they can just mooch of off google and bing for free.
I think you’d be right that the direct cost of running the crawler and index would not be the issue. But fighting SEO to keep your results decent is probably a cost that dwarfs the basic technical cost of running the crawler and index.
And you’d need a technical security team on top of things as link farms aren’t your only risk, I’m sure there are countless ways to manipulate the algorithm to put your site on top that Google probably have multiple teams working on fighting it full time.
Many of these things would likely not be a problem for a startup, though. No one is paying SEO firms big money to get into a search index no one has heard of and hardly anyone uses, so these costs probably grow exponentially over time as you become more well known.
Yeah, and on the smaller / earlier side of a theoretical search engine company, google offers their api for free. I think this is actually another one of the biggest contributors to why nobody has tried to make a new search engine with their own index. Why waste hundreds of thousands of dollars in hardware, and even more on personnel costs, when you can just have google do it for you instead?
Yes offering everything for free to prevent competition has been a surprisingly effective strategy for Google.