Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"In September 2022, we used the authoritative zone files for the top-level domains (TLDs) .com, .net, .info, and .org, together with top-1M website lists, to find a total of 255,315,270 unique names. We then queried DNS from each of five regions and recorded the set of IP addresses returned. The table below summarizes our findings:"

But of course zone files only list domainnames, not websites. More details needed on how they came up with 255M websites using top-1M website lists.

Using "top-1M" lists will bias the results because those sites are more likely to use large hosting providers and CDNs. In other words, it will exclude smaller, less popular websites using a smaller hosting providers.

"By looking at the CDF there are a few eye-watering observations:

Fewer than 10 IP addresses are needed to reach 20% of, or approximately 51 million, domains in the set;

100 IPs are enough to reach almost 50% of domains;

1000 IPs are enough to reach 60% of domains;

10,000 IPs are enough to reach 80%, or about 204 million, domains."

No surprise here. If we restrict the www to only what is popular, i.e., high-traffic, e.g., using concepts like "top-1M", then yes, the number of IP addresses we need will likely be fewer. For one because these sites all use the same handful of service providers that target high traffic websites. "Top" lists make it easier, more tidy to work with what the web actually may comprise.1 It lets "tech" companies and their service providers focus on what is commercially viable. Just ignore all those unpopular websites that cannot serve advertising. But what gets filtered out. We are prevented from knowing. Similarly, centralisation lends itself to convenience. There are obvious benefits. However, needless to say, centralisation is not appropriate in all cases.

1. Another analgous situation IMO is mobile apps. Allowing consumer to know what apps actually exist, i.e., all the millions of apps people have written, including the ones that will generate no revenue for anyone, is jettisoned in favour of "Top 10/Top20" lists or similar popularity filters. The "app store" middleman censors certain software and shows users only its view of the world's mobile app production. Searching is blunt. In the same way Google only shows the user its view of the www, rather than the actual www. Filtering can be useful, but ultimately these filtering decisions are for the benefit of the companies behind "app stores" or "web search engines". A "tech" company selling CDN services, or advertising services, is not a librarian at a university/public library. It has commercial interests. At some point "filtering" becomes "funneling", or herding.



Another analogy is the "tech" obsession with centralised "national" news instead of focusing on "local" news.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: