I'd define it as the web until the time that Facebook truly took off and conquered the hearts and minds of so many (that was the first huge shot in the losing war of the old web). For example, a key part of the old web was what used to be called the "blogosphere", and the blog's height was those final years before FB (in the ascendancy) and the first years of the FB era (descendancy).
Yeah, I'd say "old web" is early '90s up through about mid-2000s or so. Geocities, Tripod, Angelfire, lots of standalone web forums, the early blogosphere, no social media, etc.
I've just gotten used to this at this point. My first time on the web was somewhere around 1995 I think, although my first time on the Internet was earlier (it was a proxy through some ancient BBS and I'm pretty sure it was using gopher). Even though I was just a kid back then, clearly that makes me old now.
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
The plot features two American men who stumble upon Brigadoon, a mysterious Scottish village that appears for only one day every 100 years; one man soon falls in love with a young woman from Brigadoon. The show's song "Almost Like Being in Love" subsequently became a standard.
Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
Interesting blog post comparing the culture and incentives of CEOs in different countries? Completely gone, tried to search it up multiple times to no avail. It's barely 3 years old.
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
> It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
Content-addressed storage and error correcting codes mean that one can make bitrot astronomically unlikely with honestly minimal infra investment.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
There are mitigations in both software and hardware, but most consumer machines, by default, do almost none of that. No ECC RAM, no error correction in the filesystem.
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
> Webpages dying is probably one of the biggest design flaws of the original web.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
> The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration.
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded.
- The first link ever shortened was a CSS stylesheet on a WordPress blog.
- Someone shortened localhost on the second day.
- The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
I'd define it as the web until the time that Facebook truly took off and conquered the hearts and minds of so many (that was the first huge shot in the losing war of the old web). For example, a key part of the old web was what used to be called the "blogosphere", and the blog's height was those final years before FB (in the ascendancy) and the first years of the FB era (descendancy).
I would posit that "old web" could be defined as the period before Google Search became public (<~1997).
But maybe that's more a measure of my own age and perceptions rather than an accurate representation of the various eras of the internet/web...
Web Design Museum: Search Engines https://www.webdesignmuseum.org/exhibitions/search-engines-i...
2009-2014? That's not the old web. Or I'm old. Take your pick.
I’m old and it’s definitely not the old web. I consider the old web to be when PNGs were sliced with Fireworks and laid out in tables.
2009-2014 is Web 2.0, back when people still thought social media was a good idea.
Yeah, I'd say "old web" is early '90s up through about mid-2000s or so. Geocities, Tripod, Angelfire, lots of standalone web forums, the early blogosphere, no social media, etc.
Pre-social media is probably the watermark. That’s what sucked all the content out of the web and into walled-garden platforms.
I've just gotten used to this at this point. My first time on the web was somewhere around 1995 I think, although my first time on the Internet was earlier (it was a proxy through some ancient BBS and I'm pretty sure it was using gopher). Even though I was just a kid back then, clearly that makes me old now.
Speaking of gopher, I'm low-key hoping that becomes the new place for all the non-bot traffic. Gopher felt magical back in the pre-www days.
For me the old web is when people still had homepages. When those went away, the old web died.
Right? When I think of the old web I think of my Xena and X-Files fan sites on Geocities circa 1997
I was a few years late to the start (early 90s kid) but I am nostalgic for the vast amount of myfreewebs + dot.tk websites out there.
Hit counters on the front page was mandatory of course.
It was still a significantly different era than today. Let’s call it Middle Web. Maybe Late Middle Web.
+1 for this
I'd consider the facebook and mega era to be relatively new, the "old web" for me would be the one without centralization around a few giants, the era of random phpbb forums, private websites with "this site is under construction" banners and internet directories to find stuff.
God… I stood up so, so many phpBB instances.
Quite ironic that a link shortener which went offline for a decade or so is now posting about other sites not staying online.
0.mk, you had one job…
https://en.wikipedia.org/wiki/Brigadoon
It seems putting anything on the web that allows submit is screaming to get spammed these days. Is there any sort of spam filter that's worth using?
Remember the good ol' days when we all thought that everything on the web would exist for eternity and over.
I mean, it does usually, just not always in its original location.
No, that's just embarrassing stuff. That stays on the web forever.
Yep, it's Murphy's law of online content. Anything you want to reference later will be gone, with no archive copies. Anything you want deleted will be available forever.
Interesting blog post comparing the culture and incentives of CEOs in different countries? Completely gone, tried to search it up multiple times to no avail. It's barely 3 years old.
That random 20 year old video of some middle schoolers doing a flying kick and breaking a vending machine? Yeah, just popped up on my feed yesterday.
Especially since the Internet is literally a _messaging_ protocol.
Purevolume mention made me sad. I miss that community.
Am I getting old? 09-14 is not even close to the old web for me. The old web, to me, was back when people still published physical 'phone' books for websites.
Seriously. 2009 is several years after everyone was already saying “web 2.0”! That is nowhere near the “old web”.
The old web doesn't necessarily mean the oldest web. 12-17 years ago was very much an older fairly different era of the web that's worth analyzing even if it's on the younger side of the old web. I can definitively sympathize with your reaction though, it doesn't feel like that era was that long ago yet.
I vaguely recall those, but they were more for normies trying to get online. For me the old web is what I saw when I logged into my university gopher server and saw the advertisement for something called the World Wide Web which I could browse via lynx. Soon enough I got a PPP connection and then Mosaic/Netscape 1.0. However everything after javascript shipped (let alone CSS) is new new new. I'd almost go as far as saying if it doesn't have a tilde in the URL it's not old web... almost...
I’m vaguely remembering tilde username for our public html folders. Was that old Apache behavior?
And now I wonder if Apache is even still common. I spent so much time fiddling with apache configs.
"The old web" is 1993 to 2007. It's all been downhill ever since.
I was wondering what happened to the old web too
Webpages dying is probably one of the biggest design flaws of the original web.
I am not saying old content needs to be preserved forever, but so much content has factually been lost over time. Old logs from text-based MUDs for instance, even for MUDs that still exist today.
It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
We hot-linked to all those image hosts because we couldn't imagine them disappearing.
Archive.org had incredible foresight and if it didn't already exist, I'd call such a project a pipe dream.
Even Archive.org is rather limited in what it keeps. I know of a very large site that recently disappeared. archive only has part of the web html part of the site. Everything else is either gone or non accessible.
That's not my experience
> It's the great irony of digital media. Copying data is accurate to the bit and is preserved "as-is", but in practice, it requires someone to maintain servers, to care about it. To separate out what is worth preserving.
This is a bit idealized. In practice copying data is not quite accurate (especially in bulk) and bit-rot is a very real phenomenon, both in flight and in storage.
You sometimes encounter it when dealing with files from the early '00s, it's very common to discover a few of them are corrupt, even if they've only ever been copied between harddrives.
Content-addressed storage and error correcting codes mean that one can make bitrot astronomically unlikely with honestly minimal infra investment.
It's copyright that causes anything to disappear from the web IMO -- torrents never die.
EDIT: I am aware that unseeded torrents do in fact die. But it really doesn't take much to seed a whole hard drive's worth of rarely requested data -- this also detects bitrot and so corrects errors automatically if you're not the only copy.
If you are, there's ECC, as well as making another copy.
There are mitigations in both software and hardware, but most consumer machines, by default, do almost none of that. No ECC RAM, no error correction in the filesystem.
> We hot-linked to all those image hosts because we couldn't imagine them disappearing.
No, we hot-linked all those image hosts because we didn't want to pay to host it ourselves.
...and I had fun replacing images people directly linked from my server with less-- ahem-- savory images.
I enjoyed the emails I got from a couple people who were adamant I "hacked" their site because their "web developer" linked to images on my server... images that now said stuff like "I'm a loser bandwidth thief!", etc. (I never did use really nasty "shock" images-- just taunting stuff.)
Digital Data https://m.xkcd.com/1683/
There are many TinyMUD logs that were posted on Usenet, still to be found on Google Groups.
However, logging was controversial amongst mudders. It was almost always rude to log a private conversation without knowledge or consent; it was tacky to indiscriminately log while everyone was in the "hangout room" or Rec Room, as it were, and it was also bad form to post logs to Usenet or share them without redacting player names and other things.
But logging was built-in to most clients, and it was possible for server administrators to log (and hypothetically any malware-in-the-middle could log the cleartext, unencrypted TinyMUD TCP streams.) And many nefarious deeds by nasty players were exposed to the light when their logs were posted.
The technology is still in its infancy unfortunately, so there's no way the web could have been based on it, but I think content-addressing is the long-term play. If I click a link, there are some cases where I want the server to respond with a fresh response just for me (e.g. a website showing the weather). But often I just want whatever content was linked to (e.g. a webpage explaining a math content). In the latter case, it would be nice if the link had a hash of the content in it, and 3rd parties could host copies to keep the link working even if the original operator stopped existing.
That's pretty much how IPFS works.
It's pretty difficult to avoid without very significant tradeoffs, though. The closest is content-addressable peer-to-peer networks, but these still rely on someone keeping the information around, and they struggle to scale anywhere near as much.
> Webpages dying is probably one of the biggest design flaws of the original web.
I'd say it's one of the biggest design flaws of the current web, what with more and more content hidden behind paywalls, increasingly restricted WAFs, and rendered client-side via convoluted JavaScript.
Archiving and mirroring of old-style websites, delivered as static HTML, is simple and straightforward. 20 years from now, most web content from ~1996 to ~2015 will still be accessible, but much of today's web content probably won't.
Old web was kind of dumb anyway. You can put on the rose tinted glasses and feel elite about browsing some shitty site 20 years ago or enjoy the fruits of modern design.
Wow, that is a great domain!
The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration. Not counting cost for tokens to do the development.
Also:
> Reply to any 0.mk email and the message lands in a feedback queue the AI reads, triages, and acts on
Is this dangerous? What about jailbreaking AIs and having it delete everyone's account?
> The blog mentions "0.mk's revenue did not cover hosting" but then goes on to implement expensive AI integration.
The article doesn't seem to describe the cost of the AI solution. It does imply that it is lower than the cost of maintaining and supporting their service manually.
Site got hugged? Is there a torrent?
[dead]
I found an old database backup of 0.mk on a disk I had kept.
0.mk started in 2009 as a passion project built by three of us. We worked on it for a few hours each week around our regular jobs. We eventually closed it in 2014 because the revenue (hint: no revenue) could not cover hosting, development, and the constant work of fighting spam and reviewing abuse.
The recovered historical corpus contains 657,607 links. For this analysis, we followed every one of them.
Of the 655,178 links with safe, crawlable targets, 76.7% no longer returned a loading page. After removing repeated destinations, 78.7% of the 492,620 distinct crawlable URLs still did not load. So duplicate links are not creating the result.
I use “did not load” rather than “gone” deliberately. Some URLs returned 403 or 429 and may have blocked the crawler. Pages that returned 2xx or 3xx count as loading even when they now lead to parked domains, login walls, or removed-content notices.
There is one large distortion in the yearly data. A single account created 83,398 URLs pointing to one hostname in 2011. At URL level, 92.5% of that year did not load. Count each hostname once and the result becomes 61.7%, almost identical to 2010 and 2012.
A few things I did not expect:
- 835 restored links point at Facebook’s old photo CDN. None loaded. - The first link ever shortened was a CSS stylesheet on a WordPress blog. - Someone shortened localhost on the second day. - The longest stored URL is 38,753 characters and repeatedly says TRYING_THE_MAXIMUM_URL.
Most users came from one regional online community, so this is not a census of the whole web. It is a record of what that community shared between 2009 and 2014.
I brought 0.mk back to test whether AI can now handle enough development, spam filtering, abuse review, monitoring, and support to make the service sustainable where the original economics failed.
Happy to answer questions about the crawl, the old data, or the rebuild.
How did you managed to obtain that domain? Usually single digit or letter domains are “reserved”.
.mk appears to allow for it: https://marnet.mk/wp-content/uploads/2023/01/pravilnik-mk-mk...
“The name of the .mk domain consists of a minimum of 1 (one) and a maximum of 63 characters.”
ICANN might have set some rule for .com/.net/.org but it's not universal for all tlds
They’re not reserved, pretty easy to obtain if you have money, starting from less than $1k.
[flagged]
Proof? Using an AI checker isn't accurate.
The ENTIRE thing is AI generated. I'm not talking about the article. I'm talking about the entire website, the entire "product". https://0.mk/blog/zero-humans
Also it's super easy to tell by looking at it, way too many LLM-isms. No need for a AI checker tool.
> The name was registered in 2009 because it was the shortest URL possible: a zero, a dot, two letters. Seventeen years later the zero means something else.
Did the zero ever mean anything? It's still a 3 character domain regardless if the first character is a zero.
Bravo brat