Rescuing my 2006 Travel Blog

iamflimflam12 pts0 comments

Rescuing my 2006 Travel Blog | atomic14

🌈 ESP32-S3 Rainbow: ZX Spectrum Emulator Board!<br>Get it on Crowd Supply →

Rescuing my 2006 Travel Blog

View All Posts

read

Want to keep up to date with the latest posts and videos? Subscribe to the newsletter

Posts Ā·<br>Videos Ā·<br>ESP32 Ā·<br>Tools Ā·<br>Support

Ā« I Asked AI to Design 12 Dev Boards. It Told Me to Write a Script.

HELP SUPPORT MY WORK: If you're feeling flush then please stop by Patreon Or you can make a one off donation via ko-fi

In November 2006 I put everything I owned in my parents’ loft, had a great many<br>injections, and flew to Guatemala. I came back thirteen months later having<br>crossed fourteen countries between Antigua and Tierra del Fuego, and I’d written<br>the whole thing up as it happened — 182 posts typed into internet cafĆ©s, hostel<br>lobbies and the occasional boat.

That blog is still at travellingchris.blogspot.com. Most of the photos are still on<br>Flickr. Both the blogging service and Flickr still work, but Flickr is now a paid service with limits.<br>And Blogspot is now blogger (or is it the other way round?). Flickr has<br>changed hands three times and now charges for more than 1000 photos.

So I took exports from both and fed them into Claude: ā€œRebuild my old blog as a static site - here’s the blog export and here’s a bunch of photosā€. And this is what I now have:<br>travellingchris.atomic14.com.

The tech itself is pretty boring — Astro, static HTML, photos on Cloudflare<br>R2 - it should run for next to nothing. There were some interesting things along<br>the way.

Two archives to reconcile and fix

Google Takeout gave me an Atom feed with 942 entries in it: 182 posts, 203 real<br>comments, and 557 spam comments accumulated over the last couple of decades. Flickr<br>gave me back my 2,819 photos with some JSON metadata.

The two kind of match up, although I didn’t start uploading to Flickr until half<br>way through the trip.

Flickr covers April 2007 onwards — South America — with capture dates and some tags.

The Blog photos covers November 2006 to March 2007 — Central America.<br>Sadly, for that half of the trip, that’s all I have. And they are quite low resolution. I’m hoping somewhere I have a backup - but I haven’t found it yet.

Missing photos

Tracking down the photos for every posts was a bit of a mission. In all<br>there are 182 posts with 778 images in them.

Some of these turned out to be photos hotlinked from other places. These are pretty much gone forever.

Some of the earliest posts reference photos1.blogger.com, the image host Blogger used in 2006. Google<br>Takeout didn’t include those files. But the host is still up and serving files<br>nineteen years on. Claude wrote a script and pulled thirteen<br>photographs back out of it - so those photos are saved! Though again, quite low resolution.

Dodgy HTML

A fair number of the posts were broken in various ways - for example we’ve got<br>things like this:

src="http://farm3.static.flickr.com/2229/1518290776_54f431e902.jpg?v=0

Claude is pretty good at this kind of things though and fell back to just finding<br>things that looked like URLs.

After all this, we managed to get 766 photos out of the 778 image references. That’s not<br>bad for something that’s been rotting away for 19 years.

Fixing the overlap

The two archives overlap between April and June 2007, when I sometimes uploaded the<br>same photo on the blog and on Flickr. Those images were then resized by each service<br>and have different filename - so we end up with dupiclates.

Claude’s initial attempt to fix this was a bit naive and didn’t work well.

The problem is, the collection is full of similar pictures. Waterfalls and ancient ruins<br>can look almost like each other and there’s quite a few ā€œburstā€ shots. Here are three pairs my first pass<br>happily merged:

So Claude tried something else: cosine similarity between<br>mean-subtracted 32Ɨ32 greyscale thumbnails. 1,024 dimensions instead of 64 bits,<br>and blind to brightness and contrast. The numbers separate beautifully on this<br>collection:

score

Burst frames (must not merge)<br>≤ 0.9793

Genuine re-encodes (must merge)<br>≄ 0.9993

That worked pretty well and we didn’t lose any photos.

All the dates are wrong

None of the 2,819 Flickr photos has a geotag. These were all taken before<br>good mobile cameras were ubiquitus and I took them on all on a normal camera.

There is no location data anywhere in either export.

But there are 25,000 words of me saying where I was.

So the locations come from the writing. Each post is matched against a the places it mentions - including all my misspellings (my spelling is attrocious).

Photos actually embedded in a post are easy - every post has some kind of location<br>infomation in it. A photo in the Machu Picchu post is Machu Picchu, and it’s dated 1 September 2007.<br>That gives Claude 762 dated anchors scattered across the year. Using that, every remaining<br>photograph can be approximately placed by interpolating its own date against them.

Two things Claude got wrong first time:

Post dates are...

photos flickr posts blog claude things

Related Articles