<rss xmlns:source="http://source.scripting.com/" version="2.0">
  <channel>
    <title>Tim Sherratt</title>
    <link>https://updates.timsherratt.org/</link>
    <description></description>
    
    <language>en</language>
    
    <lastBuildDate>Wed, 15 Jul 2026 16:41:38 +1000</lastBuildDate>
    <item>
      <title>Honoured to be &#39;the most dedicated Labber&#39;</title>
      <link>https://updates.timsherratt.org/2026/07/15/honoured-to-be-the-most.html</link>
      <pubDate>Wed, 15 Jul 2026 16:41:38 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/07/15/honoured-to-be-the-most.html</guid>
      <description>&lt;p&gt;It can be pretty hard working on your own most of the time – unsure if what you&amp;rsquo;re doing is of any value or use. As I said in &lt;a href=&#34;https://updates.timsherratt.org/2026/06/30/the-future-of-the-past.html&#34;&gt;my keynote&lt;/a&gt; at the &lt;a href=&#34;https://www.glamlabs.io/events/glam-labs-futures-26&#34;&gt;GLAM Labs Futures conference&lt;/a&gt; in Edinburgh recently:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;the first half of last year &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;was pretty bleak&lt;/a&gt;, and left me wondering whether I should just walk away from all of it ­– go bushwalking, catch up on gardening, perhaps just be a historian again.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;So I was grateful for the opportunity to go to Edinburgh, find out what was happening in GLAM Labs around the world, connect with people I&amp;rsquo;d only met online, and rebuild my energy and enthusiasm.&lt;/p&gt;
&lt;p&gt;What I didn&amp;rsquo;t expect was to be given a special award for being &amp;lsquo;the most dedicated Labber&amp;rsquo;!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-0085-cropped.jpg&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;Photograph of a screen displaying the details of the award&#34;&gt;
&lt;p&gt;The full text of the award reads:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;We are thrilled to acknowledge and honour the incredible contributions of Dr. Tim Sherratt.&lt;/p&gt;
&lt;p&gt;For many years, Tim has been a true pioneer, operating as a visionary &amp;lsquo;one-person lab.&amp;rsquo; His GLAM Workbench Is an extraordinary repository of tools, tweaks, and thoughtful walkthroughs. It has served as sheer inspiration for anyone engaging in computational experiments with GLAM collections.&lt;/p&gt;
&lt;p&gt;This award is a testament to his immense generosity in sharing his knowledge and skills, and for always championing the spirit of &amp;lsquo;good hacking!&amp;rsquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;The award itself was hand-crafted in suitably experimental fashion by Olga Holownia.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-1022.resized.jpg&#34; width=&#34;600&#34; height=&#34;800&#34; alt=&#34;Photograph of me holding the award&#34;&gt;
&lt;p&gt;It even glows in the dark! Olga assures me it&amp;rsquo;s only slightly poisonous.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-1020.resized.jpg&#34; width=&#34;600&#34; height=&#34;800&#34; alt=&#34;Photograph of the award glowing in the dark.&#34;&gt;
&lt;p&gt;After a difficult couple of years it means a lot to be recognised by my international peers in this way, and reminds me that I&amp;rsquo;m never really working alone.&lt;/p&gt;
</description>
      <source:markdown>It can be pretty hard working on your own most of the time – unsure if what you&#39;re doing is of any value or use. As I said in [my keynote](https://updates.timsherratt.org/2026/06/30/the-future-of-the-past.html) at the [GLAM Labs Futures conference](https://www.glamlabs.io/events/glam-labs-futures-26) in Edinburgh recently:

&gt; the first half of last year [was pretty bleak](https://updates.timsherratt.org/2025/05/07/farewell-trove.html), and left me wondering whether I should just walk away from all of it ­– go bushwalking, catch up on gardening, perhaps just be a historian again.

So I was grateful for the opportunity to go to Edinburgh, find out what was happening in GLAM Labs around the world, connect with people I&#39;d only met online, and rebuild my energy and enthusiasm.

What I didn&#39;t expect was to be given a special award for being &#39;the most dedicated Labber&#39;!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-0085-cropped.jpg&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;Photograph of a screen displaying the details of the award&#34;&gt;

The full text of the award reads:

&gt; We are thrilled to acknowledge and honour the incredible contributions of Dr. Tim Sherratt.
&gt;
&gt; For many years, Tim has been a true pioneer, operating as a visionary &#39;one-person lab.&#39; His GLAM Workbench Is an extraordinary repository of tools, tweaks, and thoughtful walkthroughs. It has served as sheer inspiration for anyone engaging in computational experiments with GLAM collections.
&gt;
&gt; This award is a testament to his immense generosity in sharing his knowledge and skills, and for always championing the spirit of &#39;good hacking!&#39;

The award itself was hand-crafted in suitably experimental fashion by Olga Holownia.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-1022.resized.jpg&#34; width=&#34;600&#34; height=&#34;800&#34; alt=&#34;Photograph of me holding the award&#34;&gt;

It even glows in the dark! Olga assures me it&#39;s only slightly poisonous.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/img-1020.resized.jpg&#34; width=&#34;600&#34; height=&#34;800&#34; alt=&#34;Photograph of the award glowing in the dark.&#34;&gt;

After a difficult couple of years it means a lot to be recognised by my international peers in this way, and reminds me that I&#39;m never really working alone.
</source:markdown>
    </item>
    
    <item>
      <title>Hacking the archive – GLAM Labs workshop at the University of Edinburgh</title>
      <link>https://updates.timsherratt.org/2026/07/15/hacking-the-archive-glam-labs.html</link>
      <pubDate>Wed, 15 Jul 2026 15:53:18 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/07/15/hacking-the-archive-glam-labs.html</guid>
      <description>&lt;p&gt;On 24 June I gave a &lt;a href=&#34;https://netpreserve.org/event/workshop-hacking-the-archive/?occurrence=2026-06-24&#34;&gt;&amp;lsquo;Hacking the archive&amp;rsquo; workshop&lt;/a&gt; at the &lt;a href=&#34;https://efi.ed.ac.uk&#34;&gt;Edinburgh Futures Institute&lt;/a&gt;, University of Edinburgh. The workshop was organised by the EFI, the &lt;a href=&#34;https://www.nls.uk&#34;&gt;National Library of Scotland&lt;/a&gt;, and the &lt;a href=&#34;https://netpreserve.org&#34;&gt;IIPC&lt;/a&gt;. It preceded the &lt;a href=&#34;https://www.glamlabs.io/events/glam-labs-futures-26&#34;&gt;GLAM Labs Futures conference&lt;/a&gt;, also held at the EFI.&lt;/p&gt;
&lt;p&gt;The workshop was an expanded version of the &lt;a href=&#34;https://lab.slv.vic.gov.au/experiments/code-club/notes-hacking&#34;&gt;&amp;lsquo;hacking the library&amp;rsquo; session&lt;/a&gt; I ran for the State Library of Victoria&amp;rsquo;s Code Club in September last year. The aim of the workshop was first to give participants the confidence to poke around in websites – to play with urls, look under the hood to see how web pages are put together in the browser, and think about how they might be different. We then explored how you can change the way web pages look and work using bookmarklets and &lt;a href=&#34;https://updates.timsherratt.org/2025/07/17/glam-hacking-with-userscripts.html&#34;&gt;userscripts&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&#34;https://slides.com/wragge/glam-labs-workshop&#34;&gt;slides&lt;/a&gt; are fairly detailed, and you can probably try a number of the activities yourself if you&amp;rsquo;re interested.&lt;/p&gt;
&lt;iframe src=&#34;https://slides.com/wragge/glam-labs-workshop/embed&#34; width=&#34;100%&#34; height=&#34;500&#34; title=&#34;GLAM Labs workshop: Hacking the archive&#34; scrolling=&#34;no&#34; frameborder=&#34;0&#34; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;
&lt;p&gt;The bookmarklets and userscripts we experimented with all change aspects of the NLS website – such as inserting links or images. They&amp;rsquo;re all available &lt;a href=&#34;https://github.com/wragge/glam-labs-workshop&#34;&gt;in this GitHub repository&lt;/a&gt; and could be modified to work with other sites.&lt;/p&gt;
&lt;p&gt;The warm weather and jet lag meant I needed to be revived at various points by ice blocks (thanks Olga!), but it was a fun session, and by the end it was great to see people planning and building their own website hacks!&lt;/p&gt;
</description>
      <source:markdown>On 24 June I gave a [&#39;Hacking the archive&#39; workshop](https://netpreserve.org/event/workshop-hacking-the-archive/?occurrence=2026-06-24) at the [Edinburgh Futures Institute](https://efi.ed.ac.uk), University of Edinburgh. The workshop was organised by the EFI, the [National Library of Scotland](https://www.nls.uk), and the [IIPC](https://netpreserve.org). It preceded the [GLAM Labs Futures conference](https://www.glamlabs.io/events/glam-labs-futures-26), also held at the EFI.

The workshop was an expanded version of the [&#39;hacking the library&#39; session](https://lab.slv.vic.gov.au/experiments/code-club/notes-hacking) I ran for the State Library of Victoria&#39;s Code Club in September last year. The aim of the workshop was first to give participants the confidence to poke around in websites – to play with urls, look under the hood to see how web pages are put together in the browser, and think about how they might be different. We then explored how you can change the way web pages look and work using bookmarklets and [userscripts](https://updates.timsherratt.org/2025/07/17/glam-hacking-with-userscripts.html).

The [slides](https://slides.com/wragge/glam-labs-workshop) are fairly detailed, and you can probably try a number of the activities yourself if you&#39;re interested.

&lt;iframe src=&#34;https://slides.com/wragge/glam-labs-workshop/embed&#34; width=&#34;100%&#34; height=&#34;500&#34; title=&#34;GLAM Labs workshop: Hacking the archive&#34; scrolling=&#34;no&#34; frameborder=&#34;0&#34; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;

The bookmarklets and userscripts we experimented with all change aspects of the NLS website – such as inserting links or images. They&#39;re all available [in this GitHub repository](https://github.com/wragge/glam-labs-workshop) and could be modified to work with other sites.

The warm weather and jet lag meant I needed to be revived at various points by ice blocks (thanks Olga!), but it was a fun session, and by the end it was great to see people planning and building their own website hacks!
</source:markdown>
    </item>
    
    <item>
      <title>Some stats describing the National Archives of Australia collection</title>
      <link>https://updates.timsherratt.org/2026/07/15/some-stats-describing-the-national.html</link>
      <pubDate>Wed, 15 Jul 2026 14:20:48 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/07/15/some-stats-describing-the-national.html</guid>
      <description>&lt;p&gt;What&amp;rsquo;s actually &lt;em&gt;in&lt;/em&gt; the National Archives of Australia? And how much of that has been described, digitised, or opened to the public? I&amp;rsquo;ve made a few attempts at answering these questions over the years by harvesting summary statistics about every series in RecordSearch. I&amp;rsquo;ve now &lt;a href=&#34;https://doi.org/10.5281/zenodo.20281169&#34;&gt;saved all this data in Zenodo&lt;/a&gt;. The dataset description reads:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This dataset contains information and statistics describing record series held by the &lt;a href=&#34;https://www.naa.gov.au/&#34;&gt;National Archives of Australia&lt;/a&gt; (NAA). The data was harvested from RecordSearch, the NAA&amp;rsquo;s online database, using &lt;a href=&#34;https://glam-workbench.net/recordsearch/#harvest-details-of-all-series-in-recordsearch&#34;&gt;this notebook&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Under the &lt;a href=&#34;https://www.naa.gov.au/help-your-research/getting-started/commonwealth-record-series-crs-system&#34;&gt;Commonwealth Record Series System&lt;/a&gt; a &amp;lsquo;series&amp;rsquo; is defined as &amp;lsquo;a group of records that has resulted from the same accumulation or filing process or that has a similar format or information content&amp;rsquo;. Series are created by government agencies, and contain any number of items. The NAA uses series to organise, describe, and manage its holdings. For example, &lt;a href=&#34;https://recordsearch.naa.gov.au/scripts/AutoSearch.asp?O=S&amp;amp;Number=A1&#34;&gt;Series A1&lt;/a&gt; contains over 60,000 correspondence files created by the Department of External Affairs and its successors between 1903 and 1938.&lt;/p&gt;
&lt;p&gt;This dataset aims to provide an overview of the NAA&amp;rsquo;s holdings by compiling basic information about each series. It contains four data harvests created in late 2016, May 2021, April 2022, and May 2025. By comparing these different data files it is possible to observe changes in the overall shape of the collection, such as how many items are described, open to the public, and digitised. This should help future researchers explore the impact of digital access on historical research. For example, the 2016 data was analysed in &lt;a href=&#34;https://doi.org/10.5281/zenodo.5035855&#34;&gt;Sherratt, &amp;lsquo;Hacking Heritage: Understanding the Limits of Online Access&amp;rsquo;, 2019&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href=&#34;https://doi.org/10.5281/zenodo.20281169&#34;&gt;&lt;img src=&#34;https://zenodo.org/badge/DOI/10.5281/zenodo.20281169.svg&#34; alt=&#34;DOI&#34;&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve also created a Zenodo community for &lt;a href=&#34;https://zenodo.org/communities/naa-historical-data/records&#34;&gt;National Archives of Australia historical collection data&lt;/a&gt; which includes this dataset, as well as &lt;a href=&#34;https://doi.org/10.5281/zenodo.14769172&#34;&gt;annual harvests of &amp;lsquo;closed&amp;rsquo; files from 2016 to 2025&lt;/a&gt;, and details of &lt;a href=&#34;https://doi.org/10.5281/zenodo.14744050&#34;&gt;files digitised since 2021&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Unfortunately, due to &lt;a href=&#34;https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html&#34;&gt;changes in RecordSearch in 2025&lt;/a&gt; it is no longer possible to harvest this sort of data.&lt;/p&gt;
</description>
      <source:markdown>What&#39;s actually *in* the National Archives of Australia? And how much of that has been described, digitised, or opened to the public? I&#39;ve made a few attempts at answering these questions over the years by harvesting summary statistics about every series in RecordSearch. I&#39;ve now [saved all this data in Zenodo](https://doi.org/10.5281/zenodo.20281169). The dataset description reads:

&gt; This dataset contains information and statistics describing record series held by the [National Archives of Australia](https://www.naa.gov.au/) (NAA). The data was harvested from RecordSearch, the NAA&#39;s online database, using [this notebook](https://glam-workbench.net/recordsearch/#harvest-details-of-all-series-in-recordsearch).
&gt;
&gt; Under the [Commonwealth Record Series System](https://www.naa.gov.au/help-your-research/getting-started/commonwealth-record-series-crs-system) a &#39;series&#39; is defined as &#39;a group of records that has resulted from the same accumulation or filing process or that has a similar format or information content&#39;. Series are created by government agencies, and contain any number of items. The NAA uses series to organise, describe, and manage its holdings. For example, [Series A1](https://recordsearch.naa.gov.au/scripts/AutoSearch.asp?O=S&amp;Number=A1) contains over 60,000 correspondence files created by the Department of External Affairs and its successors between 1903 and 1938.
&gt;
&gt; This dataset aims to provide an overview of the NAA&#39;s holdings by compiling basic information about each series. It contains four data harvests created in late 2016, May 2021, April 2022, and May 2025. By comparing these different data files it is possible to observe changes in the overall shape of the collection, such as how many items are described, open to the public, and digitised. This should help future researchers explore the impact of digital access on historical research. For example, the 2016 data was analysed in [Sherratt, &#39;Hacking Heritage: Understanding the Limits of Online Access&#39;, 2019](https://doi.org/10.5281/zenodo.5035855).

[![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.20281169.svg)](https://doi.org/10.5281/zenodo.20281169)

I&#39;ve also created a Zenodo community for [National Archives of Australia historical collection data](https://zenodo.org/communities/naa-historical-data/records) which includes this dataset, as well as [annual harvests of &#39;closed&#39; files from 2016 to 2025](https://doi.org/10.5281/zenodo.14769172), and details of [files digitised since 2021](https://doi.org/10.5281/zenodo.14744050). 

Unfortunately, due to [changes in RecordSearch in 2025](https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html) it is no longer possible to harvest this sort of data.
</source:markdown>
    </item>
    
    <item>
      <title>The future of online archives in 2009</title>
      <link>https://updates.timsherratt.org/2026/07/15/the-future-of-online-archives.html</link>
      <pubDate>Wed, 15 Jul 2026 12:23:12 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/07/15/the-future-of-online-archives.html</guid>
      <description>&lt;p&gt;Way back in 2009, I was working in the web content team at the National Archives of Australia (NAA). We&amp;rsquo;d been doing some pretty interesting work on projects like &lt;a href=&#34;https://discontents.com.au/local-heroes/&#34;&gt;Mapping Our Anzacs&lt;/a&gt;, that brought together map-based finding aids and crowdsourced contributions. So when the NAA was asked to report on &amp;lsquo;emerging technologies for access&amp;rsquo; by the Standards for Public Access Working Group of the Council of Australasian Archives and Records Authorities (CAARA) the job ended up with me.&lt;/p&gt;
&lt;p&gt;In my usual way, I expanded the scope of the report until it was ridiculously over-ambitious, then had to cut it all back to try and get it finished, so I was never really satisfied with the end result. It also wasn&amp;rsquo;t a very happy time at the NAA, as the web content section was in the process of being disbanded (the usual combination of internal jealousies and fucked up management) and I was grappling with sarcoidosis. As soon as I submitted the report, I left the NAA for a placement at the National Museum of Australia and never went back.&lt;/p&gt;
&lt;p&gt;I have no idea whether the report was ever submitted to CAARA. I published it on Scribd to get it out in the world and it&amp;rsquo;s sat there ever since. However, Scribd has become increasingly enshittified, so I thought I should finally put a copy in a real, public repository. Here it is in Knowledge Commons – &lt;a href=&#34;https://doi.org/10.17613/ndffh-f5w61&#34;&gt;Emerging Technologies for the Provision of Access to Archives: Issues, Challenges and Ideas&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-07-15-12-03-07.png&#34; width=&#34;600&#34; height=&#34;728&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://doi.org/10.17613/ndffh-f5w61&#34;&gt;Read/download on Knowledge commons&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Many of the links are broken, of course, but quite a few of the issues still seem relevant, even if our optimism has faded over the years. In &lt;a href=&#34;https://updates.timsherratt.org/2026/06/30/the-future-of-the-past.html&#34;&gt;my recent GLAM Lab Futures keynote&lt;/a&gt; I suggested we should share and celebrate our own histories as a sort of counter balance to the relentless push of &amp;lsquo;progress&amp;rsquo;. So in that spirit, here&amp;rsquo;s one perspective on the future of online archives from 17 years in the past.&lt;/p&gt;
</description>
      <source:markdown>Way back in 2009, I was working in the web content team at the National Archives of Australia (NAA). We&#39;d been doing some pretty interesting work on projects like [Mapping Our Anzacs](https://discontents.com.au/local-heroes/), that brought together map-based finding aids and crowdsourced contributions. So when the NAA was asked to report on &#39;emerging technologies for access&#39; by the Standards for Public Access Working Group of the Council of Australasian Archives and Records Authorities (CAARA) the job ended up with me.

In my usual way, I expanded the scope of the report until it was ridiculously over-ambitious, then had to cut it all back to try and get it finished, so I was never really satisfied with the end result. It also wasn&#39;t a very happy time at the NAA, as the web content section was in the process of being disbanded (the usual combination of internal jealousies and fucked up management) and I was grappling with sarcoidosis. As soon as I submitted the report, I left the NAA for a placement at the National Museum of Australia and never went back.

I have no idea whether the report was ever submitted to CAARA. I published it on Scribd to get it out in the world and it&#39;s sat there ever since. However, Scribd has become increasingly enshittified, so I thought I should finally put a copy in a real, public repository. Here it is in Knowledge Commons – [Emerging Technologies for the Provision of Access to Archives: Issues, Challenges and Ideas](https://doi.org/10.17613/ndffh-f5w61).

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-07-15-12-03-07.png&#34; width=&#34;600&#34; height=&#34;728&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://doi.org/10.17613/ndffh-f5w61&#34;&gt;Read/download on Knowledge commons&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;

Many of the links are broken, of course, but quite a few of the issues still seem relevant, even if our optimism has faded over the years. In [my recent GLAM Lab Futures keynote](https://updates.timsherratt.org/2026/06/30/the-future-of-the-past.html) I suggested we should share and celebrate our own histories as a sort of counter balance to the relentless push of &#39;progress&#39;. So in that spirit, here&#39;s one perspective on the future of online archives from 17 years in the past.
</source:markdown>
    </item>
    
    <item>
      <title>The future of the past: GLAM innovation and the responsibilities of history</title>
      <link>https://updates.timsherratt.org/2026/06/30/the-future-of-the-past.html</link>
      <pubDate>Tue, 30 Jun 2026 00:08:58 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/06/30/the-future-of-the-past.html</guid>
      <description>&lt;p&gt;&lt;em&gt;Keynote presented to the &lt;a href=&#34;https://www.glamlabs.io/events/glam-labs-futures-26&#34;&gt;GLAM Labs Futures conference&lt;/a&gt;, Edinburgh, 26 June 2026. View the &lt;a href=&#34;https://slides.com/wragge/glam-labs-2026&#34;&gt;full set of slides&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;In 2012, I was lucky enough to be a Harold White Fellow at the National Library of Australia, working with data from digitised newspapers available through the Library&amp;rsquo;s innovative online service, &lt;a href=&#34;https://trove.nla.gov.au/&#34;&gt;Trove&lt;/a&gt;. I think it was the first time one of the fellowships had been awarded to a digital project.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-03.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing the Trove newspapers homepage circa 2012 and me giving the Harold White Lecture.&#34;&gt;
&lt;p&gt;I&amp;rsquo;d been playing around with Trove data for a couple of years, and had created tools to harvest and visualise searches in the digitised newspapers. There was no API back then, so everything was precariously balanced atop a series of screen scrapers. At one point, I actually pushed my scraper code to Google&amp;rsquo;s AppEngine to provide an &amp;lsquo;unofficial&amp;rsquo; API.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-04.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing some of my early visualisations of Trove data and the documentation for my &#39;unofficial&#39; API&#34;&gt;
&lt;p&gt;My fellowship project drew on some of the themes of my history PhD, which had examined ideas of progress in 20th century Australia. My plan was to use the newspapers to explore how people over the past 150 years had imagined &amp;lsquo;the future&amp;rsquo;.  I&amp;rsquo;m sure it&amp;rsquo;ll surprise no-one here to learn that I spent most of the three months of my fellowship trying to clean up OCR errors – few historical insights were forthcoming.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-05.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a series of tweets in which I&#39;m sharing the groups of TF-IDF values&#34;&gt;
&lt;p&gt;However, there was one experiment that I still enjoy. I harvested a collection of 40,000 articles that included the phrase &amp;lsquo;the future&amp;rsquo;, and grouped them by year. Then I found the words for each year that had the highest TF-IDF values – so not the most common words, but the words that were most distinctive when compared with the whole collection. I was running this process late at night, and started sharing the results over Twitter. The extracted words seemed evocative, almost poetic – so I decided to make something that people could use to craft their own odd little poems from the dataset.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-07-02-11-45-36.png&#34; width=&#34;600&#34; height=&#34;419&#34; alt=&#34;Image of The Future of the Past interface&#34;&gt;
&lt;p&gt;This was &lt;a href=&#34;https://wraggelabs.com/fotp/&#34;&gt;&amp;lsquo;The Future of the Past&amp;rsquo;&lt;/a&gt;. The interface was obviously inspired by fridge magnet poetry. You drilled down through randomly selected, TF-IDF weighted words until you reached a year. Then you dragged words around to create your poems and share them on Twitter. It was also a way of exploring the collection. Words were linked to articles, which all linked back to Trove.&lt;/p&gt;
&lt;p&gt;Over the years, the application gradually rusted, seized, and fell apart. It was running in Django and MySQL and I think I missed some updates or database migrations. When my webhost stopped supporting Python, getting it working again just seemed too hard.&lt;/p&gt;
&lt;p&gt;Last year I was doing some housekeeping, trying to bring a lot of my old apps an experiments together to reduce both the maintenance burden and my cloud hosting bills. I realised I could convert the old app to run in Flask and access its data from SQLite. Of course, Twitter had by then congealed into a slimy hellhole of neo-nazis and transphobes, so I also changed the sharing options to include Mastodon and Bluesky.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-07.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing details of my lecture include a slide from the 2012 presentation&#34;&gt;
&lt;p&gt;I thought I should also publish my fellowship lecture somewhere to provide a bit of context. The lecture has long since disappeared from the National Library&amp;rsquo;s website, but I managed to find a recording in the Internet Archive that I could transcribe and &lt;a href=&#34;https://updates.timsherratt.org/2025/06/30/mining-for-meanings.html&#34;&gt;pop into Zenodo&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html&#34;&gt;So &amp;lsquo;The Future of the Past&amp;rsquo; lives again!&lt;/a&gt; An experiment examining how the past imagined the future, itself disappeared into history, until resurrected it in the present to provide a record for the future.&lt;/p&gt;
&lt;p&gt;GLAM Labs, and GLAM innovation in general, occupy a complex position with respect to time. We explore how new technologies can be used to mobilise the past in the present. We imagine future audiences, and reconstruct past lives. We think about what&amp;rsquo;s coming next, but also what needs to be preserved.&lt;/p&gt;
&lt;p&gt;In my talk today I want to think a bit about how we navigate time.&lt;/p&gt;
&lt;hr&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-08.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Cover of the Atomic Age and Industrial Exhibition showing the atomic genie&#34;&gt;
&lt;p&gt;While researching my PhD, I found out that an Atomic Age exhibition toured Australian cities in 1947 and 1948. The exhibition included a diorama that represented the first atomic bomb test in the New Mexico desert. Emerging from the fireball, a mysterious figure loomed over the scientists – the atomic genie had been released and awaited our command. Would we use its powers to foster progress or wreak destruction?&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-09.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a crossroads image alongside images from the Atomic Age exhibition&#34;&gt;
&lt;p&gt;This choice was made even more explicit by a signpost in the middle of the exhibition. &amp;lsquo;Progress&amp;rsquo; pointed to a display of the possibilities of atomic energy in industry, while &amp;lsquo;Destruction&amp;rsquo; directed visitors to a scale model of Hiroshima with a recorded soundtrack and flashing lights for that authentic atomic annihilation experience. The idea that humankind was at some sort of &amp;lsquo;crossroads&amp;rsquo; was a common way of representing the challenges of the atomic age. We had arrived at a critical moment in history, when our decisions would determine the fate of civilisation itself.&lt;/p&gt;
&lt;p&gt;So what happened? Did we choose? This was not the first turning point that humankind had faced, nor the last. In 1966, Elizabeth Eisenstein, a historian whose major work focused on the impact of the printing press, wrote about history and our perceptions of time. She argued that linear, episodic structure of &amp;lsquo;history book time&amp;rsquo; dumps us at the opening of &amp;lsquo;the most personally significant, densely packed, fact-crowded final chapter&amp;rsquo;. The past trails off into irrelevance as we confront an unknown future full of unprecedented challenges. We are, she says, &amp;lsquo;destined always to be poised as an adult on the threshold of a new age, where previous experience offers no sure guide&amp;rsquo;.&lt;/p&gt;
&lt;p&gt;The genie is out of the bottle, there&amp;rsquo;s no turning back.&lt;/p&gt;
&lt;p&gt;The question is not whether real crises exist, but whether our perception of time helps or hinders our efforts to address them. We imagine ourselves in an eternal present where the past is closed off, and the future empty. Solutions always lay ahead.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;It took a couple of attempts to finally complete my PhD. In between, I worked for the Australian Science Archives Project (ASAP), a small, self-funded organisation attached to the University of Melbourne. Our mission was to preserve and make accessible the history of Australian science, but with limited funds for outreach we had to be a bit creative. In the early 1990s, I started converting finding aids to plain text files and loading them on to FTP and Gopher servers. Then in 1994 we took the leap to a new, exciting online platform – the web.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-10.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Archived home pages of ASAP and Bright Sparcs&#34;&gt;
&lt;p&gt;The &lt;a href=&#34;https://www.asap.unimelb.edu.au/index.html&#34;&gt;ASAP website&lt;/a&gt; was one of the first Australian history sites on the web, and certainly the first archives site in Australia. As well as newsletters and finding aids, we published a database with information about hundreds of Australian scientists and any related archival holdings – it was originally called Bright SPARCS. In the days before things like MySQL, this meant I had to figure out how to use Visual Basic to convert the Microsoft Access database into what we would now call a &amp;lsquo;static&amp;rsquo; site – lots and lots of little HTML files.&lt;/p&gt;
&lt;p&gt;After my second, successful PhD attempt, and some time as a postdoc researching the history of meteorology, I ended up back in the archives world at the National Archives of Australia (NAA). I was a member of the small web content team, and in 2008 I had an idea for a web application to accompany a new physical exhibition on Australia&amp;rsquo;s involvement in World War I.&lt;/p&gt;
&lt;p&gt;For a non-Australian audience, I feel I need to explain at this point that World War I is still a big deal in Australia. The qualities of Australia&amp;rsquo;s fighting men – the Anzacs – were mythologised and woven into a particular vision of national identity that still wields considerable political and cultural power. The exhibition I worked on was designed to highlight the digitisation of 376,000 WWI service records, funded through a special allocation from the federal government, and presented as &amp;lsquo;a gift to the nation&amp;rsquo;.&lt;/p&gt;
&lt;p&gt;The archivists who described the service records had the foresight to embed some structured data, such as places of birth and enlistment, in the file titles. So I suggested we extract the place names, geolocate them, and create a map interface for users to explore the records by location. Sounds pretty standard these days, but there was nothing quite like it at the time. We also collected photos and stories from users by setting up a &amp;lsquo;scrapbook&amp;rsquo; in Tumblr and linking it to the map interface through the Tumblr API.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-11.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of Mapping Our Anzacs and the scrapbook, as well as a photo of me introducing it to the then Prime Minister, Kevin Rudd&#34;&gt;
&lt;p&gt;The site, named &amp;lsquo;Mapping Our Anzacs&amp;rsquo; (not the choice of the development team), was popular with users who added more than 1,000 scrapbook posts in the first six months. It was also popular with politicians and bureaucrats who trumpeted it as an example of how &amp;lsquo;web 2.0&amp;rsquo; might transform government services through digital innovation and online engagement.&lt;/p&gt;
&lt;p&gt;It lasted about 6 years in its original form. In 2014, the content was rolled into a new site called &amp;lsquo;Discovering Anzacs&amp;rsquo; with some additional records. That site was suddenly decommissioned in 2023, breaking all the links that people had made to individual records, and discarding all their contributions. The &lt;a href=&#34;https://www.naa.gov.au/about-us/media-and-publications/media-releases/discovering-anzacs-website-decommissioned-making-way-innovative-new-digital-experiences&#34;&gt;media release&lt;/a&gt; announcing the change pointed people to an &amp;lsquo;archived version&amp;rsquo; in the Australian Web Archive. It was headed: &amp;lsquo;Discovering Anzacs website decommissioned, making way for innovative new digital experiences&amp;rsquo;. These new digital experiences have yet to emerge.&lt;/p&gt;
&lt;p&gt;The past is closed off and the future is empty. The media release made it seem as if the change was inevitable – technology had simply moved on. There&amp;rsquo;s no escaping the fact that long-term maintenance is hard, but there are always choices to be made. Let&amp;rsquo;s not simply shrug and point to the pace of change as a way of avoiding responsibility.&lt;/p&gt;
&lt;p&gt;&lt;a href=&#34;https://www.asap.unimelb.edu.au/bsparcs/&#34;&gt;Bright Sparcs&lt;/a&gt;, on the other hand, is still online. The original site is archived and its urls preserved. The content and identifiers have been rolled forward into the &amp;lsquo;&lt;a href=&#34;https://www.eoas.info&#34;&gt;Encyclopedia of Australian Science and Innovation&lt;/a&gt;&amp;rsquo;. After 32 years, it still works. That&amp;rsquo;s mainly due to the efforts of Gavan McCarthy, the former director of ASAP, who created the original database in the 1980s and continues to maintain it. There are always choices to be made.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;In 2009, I was &lt;a href=&#34;https://discontents.com.au/local-heroes/&#34;&gt;asked to reflect on &amp;lsquo;Mapping Our Anzacs&amp;rsquo;&lt;/a&gt; for a book edited by Kate Theimer on significance of &amp;lsquo;web 2.0&amp;rsquo; for archives and local collections. When asked what advice I&amp;rsquo;d give to an organisation venturing down this path I suggested:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Start experimenting. The technology is developing so rapidly that if you spend 12 months planning a project it’s likely to be out-of-date even before you start. New web services and data sources are becoming available every day.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Obviously, I&amp;rsquo;m still a strong believer in the value of experimentation. But I look at the sentence on the speed of change now and think it could have come from the mouth of some corporate AI shill. Quick, don&amp;rsquo;t be left behind! You can&amp;rsquo;t afford to wait! Sign up now!&lt;/p&gt;
&lt;p&gt;Maybe I&amp;rsquo;m getting old and slow, but I&amp;rsquo;m more inclined now to think about the range of timescales across which we work – about the traces we leave behind, as well as the short term impacts. If I was starting &amp;lsquo;Mapping Our Anzacs&amp;rsquo; again, I think I&amp;rsquo;d be trying to make sure all the geolocated metadata was properly versioned and saved in an open repository. Similarly, I&amp;rsquo;d create an independent backup of the scrapbook posts. I was focused on meeting the deadline and getting it to work, but I also should&amp;rsquo;ve been thinking about what happens when the institution pulls the plug.&lt;/p&gt;
&lt;p&gt;A lot of important work has been done since then on digital preservation, the value and ethics of maintenance, and planning for the death of projects. But I also wonder what the fate of our digital projects tells us about our orientation in time. For the NAA, &amp;lsquo;Mapping Our Anzacs&amp;rsquo; was a burden inherited from a near-forgotten past. For me it was an example of what you can achieve on a tiny budget by hooking together existing services and opening yourself to the public. Even after 18 years, it still seems to address the future.&lt;/p&gt;
&lt;p&gt;The work we do is embedded within its own histories – personal, institutional, technological. We find in those histories points of meaning and connection that help us make sense of where we are. I&amp;rsquo;m sure we can all point to projects or people that jolted our understanding of what was possible and sent us careening down new pathways. None of us start from scratch. For me, the period between 2007 and 2012 really helped to define what I do and why.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-14.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Screenshots from Mitchell Whitelaw&#39;s Visible Archive blog showing the Series Browser and A1 Explorer.&#39;&#34;&gt;
&lt;p&gt;In 2008, Mitchell Whitelaw was granted a fellowship by the National Archives of Australia to undertake his &lt;a href=&#34;https://visiblearchive.blogspot.com&#34;&gt;&amp;lsquo;Visible Archive&amp;rsquo; project&lt;/a&gt;. It was the first in a series of GLAM collection visualisation projects through which Mitchell developed his oft-cited concept of &amp;lsquo;generous interfaces&amp;rsquo;. I helped Mitchell wrangle some of the NAA data, and his work inspired me to look at collections as a whole, rather than as a series of individual items.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-15.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of the Zotero homepage from 2008 and some of the code from my translator&#34;&gt;
&lt;p&gt;Also in 2008, I created my first Zotero translator for the National Archives online database, RecordSearch. It made me think about what happens when we liberate collection data from web interfaces. Zotero, along with Omeka, was the product of the Center for History and New Media at George Mason University – a site that bubbled with GLAM-related enthusiasm and encouraged us to embrace the constructive power of hacking.&lt;/p&gt;
&lt;p&gt;In 2009, I visited CHNM to present &amp;lsquo;Mapping Our Anzacs&amp;rsquo; at the American Association for History and Computing conference. On the same trip I spoke at the New York Public Library, which had started pushing out a series of groundbreaking digital projects, like &amp;lsquo;Map Warper,&amp;rsquo; &amp;lsquo;Building Inspector&amp;rsquo;, and &amp;lsquo;What&amp;rsquo;s on the menu?&amp;rsquo;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-16.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of the THATCamp Canberra site as well as the list of early THATCamps&#34;&gt;
&lt;p&gt;People weren&amp;rsquo;t just experimenting with code, they were building new structures to enlarge the space and meaning of innovation. CHNM gave us THATCamp, a series of DIY unconferences that connected Digital Humanities (DH) and GLAM practitioners across institutional and disciplinary boundaries. Sick of watching events from afar, I organised THATCamp Canberra in 2010 – it was probably the most fulfilling, and exhausting, thing I&amp;rsquo;ve ever done. And if you want to know what we discussed in 2010, or in the 2011 and 2014 sequels, you can – because I&amp;rsquo;ve &lt;a href=&#34;https://thatcampcanberra.org/2011/archive-2010/index.html&#34;&gt;archived the sites&lt;/a&gt; and continue to pay the hosting bills.&lt;/p&gt;
&lt;p&gt;A few months after THATCamp Canberra, we were visited by Bethany Nowviskie, the Director of the Scholars&#39; Lab in the University of Virginia Library. Bethany challenged us all to think about the institutional, human, and political contexts of DH and GLAM innovation. In her 2011 talk, &lt;a href=&#34;https://nowviskie.org/2011/a-skunk-in-the-library/&#34;&gt;&amp;lsquo;A skunk in the library&amp;rsquo;&lt;/a&gt;, Bethany described the Scholars&#39; Lab as&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;a conscious experiment: an experiment in modeling effective relationships of research-and-development work by librarians &amp;amp; library IT both to the digital humanities as an exciting community of practice, &amp;amp; to our own future – the future of libraries within a scholarly communications ecosystem experiencing rapid reconfiguration.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-18.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;&#34;&gt;
&lt;p&gt;By 2010, I&amp;rsquo;d left the Archives and was working part-time at the National Museum of Australia, where we created our own under-the-radar, skunky GLAM Lab.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-19.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing the archived home page of the NMA Labs including a visualisation built from tiny thumbnails of collection items&#34;&gt;
&lt;p&gt;There were also new ways to play around with data. GLAM institutions had started using the Flickr Commons to share their image collections, and Flickr had an API that could be used to extract data and make new connections. One of my early experiments was the Flickr Machine Tag Challenge, which encouraged people to annotate photos with machine-readable identifiers for subjects or creators. This put me in touch others exploring the potential of Linked Open Data, and John Voss invited me to be part of the first LOD-LAM summit in San Francisco in 2011.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-20.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing archived screenshots from the Flickr Machine Tag Challenge and the LOD-LAM summit&#34;&gt;
&lt;p&gt;Communities developed, online and in-person, to share ideas and enthusiasm. In 2011, I popped across to New Zealand for my first experience of the &lt;a href=&#34;https://www.ndf.org.nz&#34;&gt;National Digital Forum&lt;/a&gt;. It was full of GLAM people doing cool digital stuff, and they were all so welcoming and generous. It felt liking coming home. The keynote speakers that year included Mitchell Whitelaw on &amp;lsquo;generous interfaces&amp;rsquo;, and Michael Lascarides on digital innovation at the NYPL.&lt;/p&gt;
&lt;p&gt;The point of these potted histories isn&amp;rsquo;t to invoke nostalgia, or suggest that some magic has been lost. There&amp;rsquo;s no lessons to be learned. History is always a conversation between past and present. What might in some respects seem to be a positive story of my growth and development, is also a catalogue of my failings – the discomfort I felt in large institutions, my tendency to self-sabotage, my impatience with administration.&lt;/p&gt;
&lt;p&gt;And not all these stories had happy endings. I created the Zotero translator for the National Archives database in my own time. When I released it, an alarmed email was circulated amongst the senior management titled &amp;lsquo;What has Tim done to Recordsearch?&amp;rsquo;. While &amp;lsquo;Mapping our Anzacs&amp;rsquo; was a great success, the web content team that created it was seen as a problem. One senior manager thought we had too many PhDs. Our positions were redefined, and our roles limited to cutting and pasting content that others had created into the content management system. We all left.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m sure many of you have similar war stories. We&amp;rsquo;ve all seen GLAM Labs come and go – projects die, initiatives falter.  But that&amp;rsquo;s all the more reason why we should remember.&lt;/p&gt;
&lt;p&gt;History provides ballast to keep us upright amidst the storms. We can draw on the strength of past achievements, reflect on failures, marshal precedents to confront new challenges. History gives us the weight and resolve to stand against the assumption of inevitability, the fetishistic power of the &amp;lsquo;new&amp;rsquo; – to ask the questions that need to be asked.&lt;/p&gt;
&lt;p&gt;When the LinkedIn bros warn that GLAM organisations are being left behind by the latest AI developments, I think about how long we&amp;rsquo;ve been working with technologies like machine learning and computer vision. Dipping again into my own history, I remember 2008, when the Powerhouse Museum started &lt;a href=&#34;https://web.archive.org/web/20080704141010/http://www.powerhousemuseum.com/dmsblog/index.php/2008/03/31/opac20-opencalais-meets-our-museum-collection-auto-tagging-and-semantic-parsing-of-collection-data&#34;&gt;using natural language processing to automatically tag collection items&lt;/a&gt;. I remember 2010, when Paul Hagon from the National Library of Australia gave a conference paper on &lt;a href=&#34;https://www.paulhagon.com/2010/03/11/everything-i-know-about-cataloguing-i-learned-from-watching-james-bond/&#34;&gt;using facial detection to explore image collections&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-23.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of the Real Face of White Australia&#39;s wall of faces&#34;&gt;
&lt;p&gt;In 2011, Paul&amp;rsquo;s work inspired Kate Bagnall and me to use facial detection to find the people inside the records of Australia&amp;rsquo;s racist migration policies and expose &lt;a href=&#34;https://www.realfaceofwhiteaustralia.net/&#34;&gt;&amp;lsquo;The Real Face of White Australia&amp;rsquo;&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-24.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of a sample of the redactions extracted from ASIO files&#34;&gt;
&lt;p&gt;A few years later, I started fiddling around with object detection to &lt;a href=&#34;https://wraggelabs.com/owebrowse/redactions/&#34;&gt;extract thousands of redactions&lt;/a&gt; from the surveillance files of Australia&amp;rsquo;s internal security organisation.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-25.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a scarf covered in ASIO redactions&#34;&gt;
&lt;p&gt;And if you want to erase yourself from history, try &lt;a href=&#34;https://updates.timsherratt.org/2021/04/21/secrets-and-lives.html&#34;&gt;wrapping yourself in one of my #redactionart scarves&lt;/a&gt;, made from 100% recycled redactions.&lt;/p&gt;
&lt;p&gt;Of course the technologies have changed, but there are continuities as well. The simplicity of turning points rarely withstands the scrutiny of history. Understanding is born from our struggle to reconcile the fact that everything is new, and yet nothing is new.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;I spent a lot of time during my PhD destroying my eyesight with microfilm readers – trawling through newspapers year by year, decade by decade.  If I&amp;rsquo;d started my research in the post-Trove era, my experience would have been very different. I wonder whether the questions I asked would have changed as well.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-26.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing some early visualisations of the number of newspaper articles in Trove&#34;&gt;
&lt;p&gt;But Trove itself isn&amp;rsquo;t fixed in time, it has its own history that runs parallel to the explorations of its users. By a sort of happy accident, some of &lt;a href=&#34;https://timsherratt.au/shed/trove/graphs/&#34;&gt;my early visualisations&lt;/a&gt; captured the state of the newspaper corpus as it was in 2011.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-27.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots from the Trove newspaper data dashboard&#34;&gt;
&lt;p&gt;I repeated the same analysis at &lt;a href=&#34;https://doi.org/10.5281/zenodo.6471544&#34;&gt;irregular intervals&lt;/a&gt; until 2022, when I set up an automated process in GitHub that captured weekly changes and &lt;a href=&#34;https://wragge.github.io/trove-newspaper-totals/&#34;&gt;displayed them on a dashboard&lt;/a&gt;. That continued until February 2025 when the National Library of Australia &lt;a href=&#34;https://updates.timsherratt.org/2025/04/11/update-on-trove-data-access.html&#34;&gt;cancelled my API access&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-28.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a chart that compares the number of newspaper articles per year in Trove in 2011 and 2022&#34;&gt;
&lt;p&gt;I&amp;rsquo;ve often used these visualisations to encourage people to &lt;a href=&#34;https://tdg.glam-workbench.net/newspapers-and-gazettes/newspaper-corpus.html&#34;&gt;think about the way the online collections are constructed&lt;/a&gt; – about how their search results are affected by things like institutional policy, legislation, funding, and technology. Sometimes I&amp;rsquo;d compare visualisations from 2011 and 2022 and ask them how their research might have been different according to their own location in time.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s been a lot of useful thinking around how we measure the value and impact of digital resources in the GLAM sector – including detailed frameworks like &lt;a href=&#34;https://pro.europeana.eu/page/impact&#34;&gt;Europeana&amp;rsquo;s Impact Playbook&lt;/a&gt; and Adrian Kingston&amp;rsquo;s Audience Impact Model. But the windows through which we observe impact are still pretty small.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m thinking of a researcher in 20 or 30 &amp;lsquo;years time who wants to understand how digital collections, like Trove, changed the practice of history – changed the types of questions we could ask about the past. They could mine the historical literature, extracting citations and analysing data use, but that only gives half of the picture. How can they examine the literature in the context of the digital collections as they were when the original research was conducted? Online collections grow as more material is digitised. Improvements in OCR make more items findable. Interface updates can affect access to the underlying data. How do we capture these sorts of changes?&lt;/p&gt;
&lt;p&gt;This history – the history of digitisation, metadata enrichment, interface design, prototype construction, dataset documentation, tool development – asserts the value of what we do. It matters. It changes things. Our projects might disappear as institutional priorities shift, but they are not disposable. They should not be forgotten.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-29.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of the datasets available in the Trove historical data community on Zenodo&#34;&gt;
&lt;p&gt;In a gesture towards that hypothetical future researcher, I&amp;rsquo;ve assembled an &lt;a href=&#34;https://zenodo-rdm.web.cern.ch/communities/trove-historical-data/&#34;&gt;idiosyncratic collection of snapshots and datasets&lt;/a&gt; in Zenodo. They include things like the &lt;a href=&#34;https://doi.org/10.5281/zenodo.11496377&#34;&gt;2,495,958 public tags added to 10,403,650 resources in Trove from 2008 to 2024&lt;/a&gt;, &lt;a href=&#34;https://doi.org/10.5281/zenodo.13761534&#34;&gt;lists of non-English newspapers in Trove&lt;/a&gt;, and &lt;a href=&#34;https://doi.org/10.5281/zenodo.13761546&#34;&gt;the number of OCR corrections&lt;/a&gt; by year, article category, and newspaper title.&lt;/p&gt;
&lt;p&gt;Perhaps my favourite set of collection snapshots comes from the National Archives of Australia&amp;rsquo;s RecordSearch database.  The records of Australia&amp;rsquo;s federal government are supposed to available to the public after 20 years. However, some are withheld for reasons like national security and privacy. Each year, the NAA makes a big performance out of revealing newly-released cabinet records, which are duly reported by the media on 1 January. I thought it was only fair that the public should also see the list of records that were currently &lt;em&gt;closed&lt;/em&gt; to public access. &lt;a href=&#34;https://doi.org/10.5281/zenodo.14769172&#34;&gt;So every New Year&amp;rsquo;s Day for ten years, I harvested details of the files we weren&amp;rsquo;t allowed to see.&lt;/a&gt; I only stopped because the NAA &lt;a href=&#34;https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html&#34;&gt;introduced anti-bot measures that blocked my scraper script&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-30.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots  ofdataset of closed files in Zenodo and article titled &#39;Withheld pending advice&#39;&#34;&gt;
&lt;p&gt;Beyond a little sly subversion, the data on closed files in the NAA is useful because it helps to document &lt;a href=&#34;https://insidestory.org.au/withheld-pending-advice/&#34;&gt;the workings of the access examination system&lt;/a&gt; – a point of much pain for researchers. Most of my work over the past 30 years has, in one way or another, explored the meaning of &amp;lsquo;access&amp;rsquo; – how it is constructed, how that changes, and what it means for people using GLAM collections.&lt;/p&gt;
&lt;hr&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-31.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing newspaper article headed &#39;Phyllis in Atomic wonderland&#39;&#34;&gt;
&lt;p&gt;In January 1948, 13 year old Phyllis Nichols stood at the crossroads. She was visiting the Atomic Age Exhibition in Melbourne and according to the &lt;em&gt;Sun&lt;/em&gt; newspaper:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;She had covered the path of destruction and she turned with hope to the road to progress.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Such a weighty decision for a 13 year old. I must admit, there was a point where I seriously considered turning my thesis into a work of fiction focused on Phyllis&amp;rsquo;s adventures in atomic wonderland.&lt;/p&gt;
&lt;p&gt;Phyllis chose well. But the crossroads metaphor was never really about choice. No-one was expected to pursue the path to nuclear annihilation. The crossroads demanded obedience to a specific vision of the future. It&amp;rsquo;s this way &amp;hellip;or else.&lt;/p&gt;
&lt;p&gt;&amp;lsquo;Mapping Our Anzacs&amp;rsquo; was created at a time of optimism, when it was thought that web technologies would open up the workings of government to new forms of public participation and transparency. But the dreams of &amp;lsquo;government 2.0&amp;rsquo; have faded, as information becomes ever more tightly controlled. In the GLAM sector, APIs have come and gone. Datasets created for long past hack events linger without updates, almost forgotten. New defensive measures aimed at taming the onslaught of AI scraper bots have imposed extra limits on access. Meanwhile a handful of tech oligarchs tell us what our future will be. This is the reality of progress.&lt;/p&gt;
&lt;p&gt;The work of GLAM Labs, of GLAM innovation, has always been focused on expanding the realm of the possible – encouraging people to see differently, to think differently. This work struggles constantly with the many meanings of &amp;lsquo;access&amp;rsquo; – what use is data without good documentation, without permissive licences, without tools for analysis, without the skills of confidence to use those tools. It was this sort of struggle that motivated the &lt;a href=&#34;https://glam-workbench.net/&#34;&gt;GLAM Workbench&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/557c57b82c.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing home page of the GLAM Workbench&#34;&gt;
&lt;p&gt;I recently wrote a &lt;a href=&#34;https://updates.timsherratt.org/2025/06/05/glam-workbench-preprint-for-building.html&#34;&gt;potted introduction to the GLAM Workbench&lt;/a&gt; for a forthcoming publication on tool-making in the digital humanities. I won&amp;rsquo;t read it all out, but I think it gives a pretty good overview of where things are.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-33.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Screenshot showing article in Zenodo&#34;&gt;
&lt;p&gt;I suppose I want to emphasise though that the aim of the GLAM Workbench has always been to document possibilities – to expose researchers to the richness of GLAM data, to the new types of questions they can ask, and to the methods that are available to connect everything up.&lt;/p&gt;
&lt;p&gt;In a world that erects multiple barriers of expertise, ownership, participation, and authority, there&amp;rsquo;s power in simply knowing what&amp;rsquo;s possible.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Back in 2010, at the first THATCamp Canberra, someone thanked me and said &amp;lsquo;I&amp;rsquo;ve found my people&amp;rsquo;. That sense of belonging was always what made the National Digital Forum in New Zealand so special. I&amp;rsquo;m not great at organisations – meetings make me anxious, and my email is a bin fire – but I do draw a lot of strength from the passions of like-minded people.&lt;/p&gt;
&lt;p&gt;I don&amp;rsquo;t really know what the future of the GLAM Workbench will be. I&amp;rsquo;d like to be confident and optimistic, but the first half of last year &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;was pretty bleak&lt;/a&gt;, and left me wondering whether I should just walk away from all of it ­– go bushwalking, catch up on gardening, perhaps just be a historian again. I was saved by the GLAM Labs community. First of all, by the fabulous folk at the State Library of Victoria&amp;rsquo;s LAB, particularly Paula Bray and Sotirios Alpanis. They gave me what I needed – fun data and wicked challenges. For a few months, I was back in my happy place, &lt;a href=&#34;https://slv.wraggelabs.com/&#34;&gt;creating new pathways through the SLV&amp;rsquo;s place-based collections&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-34.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide with screenshots showing the page that collects all the outputs from my SLV residency and an example of viewing photos by street in the CUA browser&#34;&gt;
&lt;p&gt;The second boost was, of course, the invitation to be here today – to find out what&amp;rsquo;s happening in labs around the world, to finally meet people I&amp;rsquo;ve known online for years, to join in the excitement and, yes, to share the disappointments.&lt;/p&gt;
&lt;p&gt;Perhaps the GLAM Workbench has done it&amp;rsquo;s job. I think it&amp;rsquo;s helped give people the confidence to dip a toe in the world of collections as data. It&amp;rsquo;s also provided a useful model for GLAM organisations seeking to encourage new types of research. Perhaps its main value was always as an intervention – an invocation of possibilities that filled a particular gap at a particular moment in time. I&amp;rsquo;ve always tried to wrap the GLAM Workbench in layers of documentation to enable any lasting value to be extracted as needed. Perhaps my focus should be to make sure that documentation is complete.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-35.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide with screenshot from the GLAM Workbench describing how organisations can create their own sections/repositories&#34;&gt;
&lt;p&gt;But it&amp;rsquo;s not a choice for me alone. The GLAM Workbench has &lt;a href=&#34;https://glam-workbench.net/get-involved/developing-repositories/&#34;&gt;always welcomed contributions&lt;/a&gt;, so if you&amp;rsquo;d like to carve out your own spaces, create your own sections, let me know! Perhaps the future of the GLAM Workbench will be shaped by the hands of others.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s also a lot of cool GLAM data out there to play around with, and it really doesn&amp;rsquo;t take much to get me excited about it. So perhaps I&amp;rsquo;ll just continue to follow my enthusiasms and see where that takes things.&lt;/p&gt;
&lt;p&gt;All of these possible futures are good. The choices aren&amp;rsquo;t fixed &amp;ndash; they leave the conversation with history open and constructive. I&amp;rsquo;m happy with that.&lt;/p&gt;
&lt;p&gt;So greetings, thanks, and solidarity to all GLAM Labbers, past, present, and future. Despite all the setbacks and frustrations, your work matters. Take time to remember, to enjoy, and to celebrate.&lt;/p&gt;
</description>
      <source:markdown>*Keynote presented to the [GLAM Labs Futures conference](https://www.glamlabs.io/events/glam-labs-futures-26), Edinburgh, 26 June 2026. View the [full set of slides](https://slides.com/wragge/glam-labs-2026).*

In 2012, I was lucky enough to be a Harold White Fellow at the National Library of Australia, working with data from digitised newspapers available through the Library&#39;s innovative online service, [Trove](https://trove.nla.gov.au/). I think it was the first time one of the fellowships had been awarded to a digital project.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-03.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing the Trove newspapers homepage circa 2012 and me giving the Harold White Lecture.&#34;&gt;

I&#39;d been playing around with Trove data for a couple of years, and had created tools to harvest and visualise searches in the digitised newspapers. There was no API back then, so everything was precariously balanced atop a series of screen scrapers. At one point, I actually pushed my scraper code to Google&#39;s AppEngine to provide an &#39;unofficial&#39; API.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-04.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing some of my early visualisations of Trove data and the documentation for my &#39;unofficial&#39; API&#34;&gt;

My fellowship project drew on some of the themes of my history PhD, which had examined ideas of progress in 20th century Australia. My plan was to use the newspapers to explore how people over the past 150 years had imagined &#39;the future&#39;.  I&#39;m sure it&#39;ll surprise no-one here to learn that I spent most of the three months of my fellowship trying to clean up OCR errors – few historical insights were forthcoming.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-05.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a series of tweets in which I&#39;m sharing the groups of TF-IDF values&#34;&gt;

However, there was one experiment that I still enjoy. I harvested a collection of 40,000 articles that included the phrase &#39;the future&#39;, and grouped them by year. Then I found the words for each year that had the highest TF-IDF values – so not the most common words, but the words that were most distinctive when compared with the whole collection. I was running this process late at night, and started sharing the results over Twitter. The extracted words seemed evocative, almost poetic – so I decided to make something that people could use to craft their own odd little poems from the dataset.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-07-02-11-45-36.png&#34; width=&#34;600&#34; height=&#34;419&#34; alt=&#34;Image of The Future of the Past interface&#34;&gt;

This was [&#39;The Future of the Past&#39;](https://wraggelabs.com/fotp/). The interface was obviously inspired by fridge magnet poetry. You drilled down through randomly selected, TF-IDF weighted words until you reached a year. Then you dragged words around to create your poems and share them on Twitter. It was also a way of exploring the collection. Words were linked to articles, which all linked back to Trove.

Over the years, the application gradually rusted, seized, and fell apart. It was running in Django and MySQL and I think I missed some updates or database migrations. When my webhost stopped supporting Python, getting it working again just seemed too hard.

Last year I was doing some housekeeping, trying to bring a lot of my old apps an experiments together to reduce both the maintenance burden and my cloud hosting bills. I realised I could convert the old app to run in Flask and access its data from SQLite. Of course, Twitter had by then congealed into a slimy hellhole of neo-nazis and transphobes, so I also changed the sharing options to include Mastodon and Bluesky.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-07.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing details of my lecture include a slide from the 2012 presentation&#34;&gt;

I thought I should also publish my fellowship lecture somewhere to provide a bit of context. The lecture has long since disappeared from the National Library&#39;s website, but I managed to find a recording in the Internet Archive that I could transcribe and [pop into Zenodo](https://updates.timsherratt.org/2025/06/30/mining-for-meanings.html).

[So &#39;The Future of the Past&#39; lives again!](https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html) An experiment examining how the past imagined the future, itself disappeared into history, until resurrected it in the present to provide a record for the future. 

GLAM Labs, and GLAM innovation in general, occupy a complex position with respect to time. We explore how new technologies can be used to mobilise the past in the present. We imagine future audiences, and reconstruct past lives. We think about what&#39;s coming next, but also what needs to be preserved.

In my talk today I want to think a bit about how we navigate time.

----

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-08.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Cover of the Atomic Age and Industrial Exhibition showing the atomic genie&#34;&gt;

While researching my PhD, I found out that an Atomic Age exhibition toured Australian cities in 1947 and 1948. The exhibition included a diorama that represented the first atomic bomb test in the New Mexico desert. Emerging from the fireball, a mysterious figure loomed over the scientists – the atomic genie had been released and awaited our command. Would we use its powers to foster progress or wreak destruction? 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-09.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a crossroads image alongside images from the Atomic Age exhibition&#34;&gt;

This choice was made even more explicit by a signpost in the middle of the exhibition. &#39;Progress&#39; pointed to a display of the possibilities of atomic energy in industry, while &#39;Destruction&#39; directed visitors to a scale model of Hiroshima with a recorded soundtrack and flashing lights for that authentic atomic annihilation experience. The idea that humankind was at some sort of &#39;crossroads&#39; was a common way of representing the challenges of the atomic age. We had arrived at a critical moment in history, when our decisions would determine the fate of civilisation itself.

So what happened? Did we choose? This was not the first turning point that humankind had faced, nor the last. In 1966, Elizabeth Eisenstein, a historian whose major work focused on the impact of the printing press, wrote about history and our perceptions of time. She argued that linear, episodic structure of &#39;history book time&#39; dumps us at the opening of &#39;the most personally significant, densely packed, fact-crowded final chapter&#39;. The past trails off into irrelevance as we confront an unknown future full of unprecedented challenges. We are, she says, &#39;destined always to be poised as an adult on the threshold of a new age, where previous experience offers no sure guide&#39;. 

The genie is out of the bottle, there&#39;s no turning back. 

The question is not whether real crises exist, but whether our perception of time helps or hinders our efforts to address them. We imagine ourselves in an eternal present where the past is closed off, and the future empty. Solutions always lay ahead.

----

It took a couple of attempts to finally complete my PhD. In between, I worked for the Australian Science Archives Project (ASAP), a small, self-funded organisation attached to the University of Melbourne. Our mission was to preserve and make accessible the history of Australian science, but with limited funds for outreach we had to be a bit creative. In the early 1990s, I started converting finding aids to plain text files and loading them on to FTP and Gopher servers. Then in 1994 we took the leap to a new, exciting online platform – the web.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-10.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Archived home pages of ASAP and Bright Sparcs&#34;&gt;

The [ASAP website](https://www.asap.unimelb.edu.au/index.html) was one of the first Australian history sites on the web, and certainly the first archives site in Australia. As well as newsletters and finding aids, we published a database with information about hundreds of Australian scientists and any related archival holdings – it was originally called Bright SPARCS. In the days before things like MySQL, this meant I had to figure out how to use Visual Basic to convert the Microsoft Access database into what we would now call a &#39;static&#39; site – lots and lots of little HTML files.

After my second, successful PhD attempt, and some time as a postdoc researching the history of meteorology, I ended up back in the archives world at the National Archives of Australia (NAA). I was a member of the small web content team, and in 2008 I had an idea for a web application to accompany a new physical exhibition on Australia&#39;s involvement in World War I.

For a non-Australian audience, I feel I need to explain at this point that World War I is still a big deal in Australia. The qualities of Australia&#39;s fighting men – the Anzacs – were mythologised and woven into a particular vision of national identity that still wields considerable political and cultural power. The exhibition I worked on was designed to highlight the digitisation of 376,000 WWI service records, funded through a special allocation from the federal government, and presented as &#39;a gift to the nation&#39;.

The archivists who described the service records had the foresight to embed some structured data, such as places of birth and enlistment, in the file titles. So I suggested we extract the place names, geolocate them, and create a map interface for users to explore the records by location. Sounds pretty standard these days, but there was nothing quite like it at the time. We also collected photos and stories from users by setting up a &#39;scrapbook&#39; in Tumblr and linking it to the map interface through the Tumblr API. 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-11.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of Mapping Our Anzacs and the scrapbook, as well as a photo of me introducing it to the then Prime Minister, Kevin Rudd&#34;&gt;

The site, named &#39;Mapping Our Anzacs&#39; (not the choice of the development team), was popular with users who added more than 1,000 scrapbook posts in the first six months. It was also popular with politicians and bureaucrats who trumpeted it as an example of how &#39;web 2.0&#39; might transform government services through digital innovation and online engagement.

It lasted about 6 years in its original form. In 2014, the content was rolled into a new site called &#39;Discovering Anzacs&#39; with some additional records. That site was suddenly decommissioned in 2023, breaking all the links that people had made to individual records, and discarding all their contributions. The [media release](https://www.naa.gov.au/about-us/media-and-publications/media-releases/discovering-anzacs-website-decommissioned-making-way-innovative-new-digital-experiences) announcing the change pointed people to an &#39;archived version&#39; in the Australian Web Archive. It was headed: &#39;Discovering Anzacs website decommissioned, making way for innovative new digital experiences&#39;. These new digital experiences have yet to emerge.

The past is closed off and the future is empty. The media release made it seem as if the change was inevitable – technology had simply moved on. There&#39;s no escaping the fact that long-term maintenance is hard, but there are always choices to be made. Let&#39;s not simply shrug and point to the pace of change as a way of avoiding responsibility.

[Bright Sparcs](https://www.asap.unimelb.edu.au/bsparcs/), on the other hand, is still online. The original site is archived and its urls preserved. The content and identifiers have been rolled forward into the &#39;[Encyclopedia of Australian Science and Innovation](https://www.eoas.info)&#39;. After 32 years, it still works. That&#39;s mainly due to the efforts of Gavan McCarthy, the former director of ASAP, who created the original database in the 1980s and continues to maintain it. There are always choices to be made.

----

In 2009, I was [asked to reflect on &#39;Mapping Our Anzacs&#39;](https://discontents.com.au/local-heroes/) for a book edited by Kate Theimer on significance of &#39;web 2.0&#39; for archives and local collections. When asked what advice I&#39;d give to an organisation venturing down this path I suggested:

&gt; Start experimenting. The technology is developing so rapidly that if you spend 12 months planning a project it’s likely to be out-of-date even before you start. New web services and data sources are becoming available every day.

Obviously, I&#39;m still a strong believer in the value of experimentation. But I look at the sentence on the speed of change now and think it could have come from the mouth of some corporate AI shill. Quick, don&#39;t be left behind! You can&#39;t afford to wait! Sign up now!

Maybe I&#39;m getting old and slow, but I&#39;m more inclined now to think about the range of timescales across which we work – about the traces we leave behind, as well as the short term impacts. If I was starting &#39;Mapping Our Anzacs&#39; again, I think I&#39;d be trying to make sure all the geolocated metadata was properly versioned and saved in an open repository. Similarly, I&#39;d create an independent backup of the scrapbook posts. I was focused on meeting the deadline and getting it to work, but I also should&#39;ve been thinking about what happens when the institution pulls the plug.

A lot of important work has been done since then on digital preservation, the value and ethics of maintenance, and planning for the death of projects. But I also wonder what the fate of our digital projects tells us about our orientation in time. For the NAA, &#39;Mapping Our Anzacs&#39; was a burden inherited from a near-forgotten past. For me it was an example of what you can achieve on a tiny budget by hooking together existing services and opening yourself to the public. Even after 18 years, it still seems to address the future.

The work we do is embedded within its own histories – personal, institutional, technological. We find in those histories points of meaning and connection that help us make sense of where we are. I&#39;m sure we can all point to projects or people that jolted our understanding of what was possible and sent us careening down new pathways. None of us start from scratch. For me, the period between 2007 and 2012 really helped to define what I do and why.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-14.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Screenshots from Mitchell Whitelaw&#39;s Visible Archive blog showing the Series Browser and A1 Explorer.&#39;&#34;&gt;

In 2008, Mitchell Whitelaw was granted a fellowship by the National Archives of Australia to undertake his [&#39;Visible Archive&#39; project](https://visiblearchive.blogspot.com). It was the first in a series of GLAM collection visualisation projects through which Mitchell developed his oft-cited concept of &#39;generous interfaces&#39;. I helped Mitchell wrangle some of the NAA data, and his work inspired me to look at collections as a whole, rather than as a series of individual items.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-15.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of the Zotero homepage from 2008 and some of the code from my translator&#34;&gt;

Also in 2008, I created my first Zotero translator for the National Archives online database, RecordSearch. It made me think about what happens when we liberate collection data from web interfaces. Zotero, along with Omeka, was the product of the Center for History and New Media at George Mason University – a site that bubbled with GLAM-related enthusiasm and encouraged us to embrace the constructive power of hacking.

In 2009, I visited CHNM to present &#39;Mapping Our Anzacs&#39; at the American Association for History and Computing conference. On the same trip I spoke at the New York Public Library, which had started pushing out a series of groundbreaking digital projects, like &#39;Map Warper,&#39; &#39;Building Inspector&#39;, and &#39;What&#39;s on the menu?&#39;.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-16.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of the THATCamp Canberra site as well as the list of early THATCamps&#34;&gt;

People weren&#39;t just experimenting with code, they were building new structures to enlarge the space and meaning of innovation. CHNM gave us THATCamp, a series of DIY unconferences that connected Digital Humanities (DH) and GLAM practitioners across institutional and disciplinary boundaries. Sick of watching events from afar, I organised THATCamp Canberra in 2010 – it was probably the most fulfilling, and exhausting, thing I&#39;ve ever done. And if you want to know what we discussed in 2010, or in the 2011 and 2014 sequels, you can – because I&#39;ve [archived the sites](https://thatcampcanberra.org/2011/archive-2010/index.html) and continue to pay the hosting bills.

A few months after THATCamp Canberra, we were visited by Bethany Nowviskie, the Director of the Scholars&#39; Lab in the University of Virginia Library. Bethany challenged us all to think about the institutional, human, and political contexts of DH and GLAM innovation. In her 2011 talk, [&#39;A skunk in the library&#39;](https://nowviskie.org/2011/a-skunk-in-the-library/), Bethany described the Scholars&#39; Lab as

&gt;  a conscious experiment: an experiment in modeling effective relationships of research-and-development work by librarians &amp; library IT both to the digital humanities as an exciting community of practice, &amp; to our own future – the future of libraries within a scholarly communications ecosystem experiencing rapid reconfiguration.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-18.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;&#34;&gt;

By 2010, I&#39;d left the Archives and was working part-time at the National Museum of Australia, where we created our own under-the-radar, skunky GLAM Lab.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-19.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing the archived home page of the NMA Labs including a visualisation built from tiny thumbnails of collection items&#34;&gt;

There were also new ways to play around with data. GLAM institutions had started using the Flickr Commons to share their image collections, and Flickr had an API that could be used to extract data and make new connections. One of my early experiments was the Flickr Machine Tag Challenge, which encouraged people to annotate photos with machine-readable identifiers for subjects or creators. This put me in touch others exploring the potential of Linked Open Data, and John Voss invited me to be part of the first LOD-LAM summit in San Francisco in 2011. 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-20.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing archived screenshots from the Flickr Machine Tag Challenge and the LOD-LAM summit&#34;&gt;

Communities developed, online and in-person, to share ideas and enthusiasm. In 2011, I popped across to New Zealand for my first experience of the [National Digital Forum](https://www.ndf.org.nz). It was full of GLAM people doing cool digital stuff, and they were all so welcoming and generous. It felt liking coming home. The keynote speakers that year included Mitchell Whitelaw on &#39;generous interfaces&#39;, and Michael Lascarides on digital innovation at the NYPL.

The point of these potted histories isn&#39;t to invoke nostalgia, or suggest that some magic has been lost. There&#39;s no lessons to be learned. History is always a conversation between past and present. What might in some respects seem to be a positive story of my growth and development, is also a catalogue of my failings – the discomfort I felt in large institutions, my tendency to self-sabotage, my impatience with administration.

And not all these stories had happy endings. I created the Zotero translator for the National Archives database in my own time. When I released it, an alarmed email was circulated amongst the senior management titled &#39;What has Tim done to Recordsearch?&#39;. While &#39;Mapping our Anzacs&#39; was a great success, the web content team that created it was seen as a problem. One senior manager thought we had too many PhDs. Our positions were redefined, and our roles limited to cutting and pasting content that others had created into the content management system. We all left.

I&#39;m sure many of you have similar war stories. We&#39;ve all seen GLAM Labs come and go – projects die, initiatives falter.  But that&#39;s all the more reason why we should remember.

History provides ballast to keep us upright amidst the storms. We can draw on the strength of past achievements, reflect on failures, marshal precedents to confront new challenges. History gives us the weight and resolve to stand against the assumption of inevitability, the fetishistic power of the &#39;new&#39; – to ask the questions that need to be asked. 

When the LinkedIn bros warn that GLAM organisations are being left behind by the latest AI developments, I think about how long we&#39;ve been working with technologies like machine learning and computer vision. Dipping again into my own history, I remember 2008, when the Powerhouse Museum started [using natural language processing to automatically tag collection items](https://web.archive.org/web/20080704141010/http://www.powerhousemuseum.com/dmsblog/index.php/2008/03/31/opac20-opencalais-meets-our-museum-collection-auto-tagging-and-semantic-parsing-of-collection-data). I remember 2010, when Paul Hagon from the National Library of Australia gave a conference paper on [using facial detection to explore image collections](https://www.paulhagon.com/2010/03/11/everything-i-know-about-cataloguing-i-learned-from-watching-james-bond/). 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-23.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of the Real Face of White Australia&#39;s wall of faces&#34;&gt;

In 2011, Paul&#39;s work inspired Kate Bagnall and me to use facial detection to find the people inside the records of Australia&#39;s racist migration policies and expose [&#39;The Real Face of White Australia&#39;](https://www.realfaceofwhiteaustralia.net/).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-24.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a screenshot of a sample of the redactions extracted from ASIO files&#34;&gt;

A few years later, I started fiddling around with object detection to [extract thousands of redactions](https://wraggelabs.com/owebrowse/redactions/) from the surveillance files of Australia&#39;s internal security organisation.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-25.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a scarf covered in ASIO redactions&#34;&gt;

And if you want to erase yourself from history, try [wrapping yourself in one of my #redactionart scarves](https://updates.timsherratt.org/2021/04/21/secrets-and-lives.html), made from 100% recycled redactions.

Of course the technologies have changed, but there are continuities as well. The simplicity of turning points rarely withstands the scrutiny of history. Understanding is born from our struggle to reconcile the fact that everything is new, and yet nothing is new.

----

I spent a lot of time during my PhD destroying my eyesight with microfilm readers – trawling through newspapers year by year, decade by decade.  If I&#39;d started my research in the post-Trove era, my experience would have been very different. I wonder whether the questions I asked would have changed as well.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-26.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing some early visualisations of the number of newspaper articles in Trove&#34;&gt;

But Trove itself isn&#39;t fixed in time, it has its own history that runs parallel to the explorations of its users. By a sort of happy accident, some of [my early visualisations](https://timsherratt.au/shed/trove/graphs/) captured the state of the newspaper corpus as it was in 2011.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-27.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots from the Trove newspaper data dashboard&#34;&gt;

I repeated the same analysis at [irregular intervals](https://doi.org/10.5281/zenodo.6471544) until 2022, when I set up an automated process in GitHub that captured weekly changes and [displayed them on a dashboard](https://wragge.github.io/trove-newspaper-totals/). That continued until February 2025 when the National Library of Australia [cancelled my API access](https://updates.timsherratt.org/2025/04/11/update-on-trove-data-access.html).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-28.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing a chart that compares the number of newspaper articles per year in Trove in 2011 and 2022&#34;&gt;

I&#39;ve often used these visualisations to encourage people to [think about the way the online collections are constructed](https://tdg.glam-workbench.net/newspapers-and-gazettes/newspaper-corpus.html) – about how their search results are affected by things like institutional policy, legislation, funding, and technology. Sometimes I&#39;d compare visualisations from 2011 and 2022 and ask them how their research might have been different according to their own location in time.

There&#39;s been a lot of useful thinking around how we measure the value and impact of digital resources in the GLAM sector – including detailed frameworks like [Europeana&#39;s Impact Playbook](https://pro.europeana.eu/page/impact) and Adrian Kingston&#39;s Audience Impact Model. But the windows through which we observe impact are still pretty small.

I&#39;m thinking of a researcher in 20 or 30 &#39;years time who wants to understand how digital collections, like Trove, changed the practice of history – changed the types of questions we could ask about the past. They could mine the historical literature, extracting citations and analysing data use, but that only gives half of the picture. How can they examine the literature in the context of the digital collections as they were when the original research was conducted? Online collections grow as more material is digitised. Improvements in OCR make more items findable. Interface updates can affect access to the underlying data. How do we capture these sorts of changes?

This history – the history of digitisation, metadata enrichment, interface design, prototype construction, dataset documentation, tool development – asserts the value of what we do. It matters. It changes things. Our projects might disappear as institutional priorities shift, but they are not disposable. They should not be forgotten.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-29.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots of the datasets available in the Trove historical data community on Zenodo&#34;&gt;

In a gesture towards that hypothetical future researcher, I&#39;ve assembled an [idiosyncratic collection of snapshots and datasets](https://zenodo-rdm.web.cern.ch/communities/trove-historical-data/) in Zenodo. They include things like the [2,495,958 public tags added to 10,403,650 resources in Trove from 2008 to 2024](https://doi.org/10.5281/zenodo.11496377), [lists of non-English newspapers in Trove](https://doi.org/10.5281/zenodo.13761534), and [the number of OCR corrections](https://doi.org/10.5281/zenodo.13761546) by year, article category, and newspaper title.

Perhaps my favourite set of collection snapshots comes from the National Archives of Australia&#39;s RecordSearch database.  The records of Australia&#39;s federal government are supposed to available to the public after 20 years. However, some are withheld for reasons like national security and privacy. Each year, the NAA makes a big performance out of revealing newly-released cabinet records, which are duly reported by the media on 1 January. I thought it was only fair that the public should also see the list of records that were currently *closed* to public access. [So every New Year&#39;s Day for ten years, I harvested details of the files we weren&#39;t allowed to see.](https://doi.org/10.5281/zenodo.14769172) I only stopped because the NAA [introduced anti-bot measures that blocked my scraper script](https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-30.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing screenshots  ofdataset of closed files in Zenodo and article titled &#39;Withheld pending advice&#39;&#34;&gt;

Beyond a little sly subversion, the data on closed files in the NAA is useful because it helps to document [the workings of the access examination system](https://insidestory.org.au/withheld-pending-advice/) – a point of much pain for researchers. Most of my work over the past 30 years has, in one way or another, explored the meaning of &#39;access&#39; – how it is constructed, how that changes, and what it means for people using GLAM collections.

----

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-31.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing newspaper article headed &#39;Phyllis in Atomic wonderland&#39;&#34;&gt;

In January 1948, 13 year old Phyllis Nichols stood at the crossroads. She was visiting the Atomic Age Exhibition in Melbourne and according to the *Sun* newspaper:

&gt; She had covered the path of destruction and she turned with hope to the road to progress.

Such a weighty decision for a 13 year old. I must admit, there was a point where I seriously considered turning my thesis into a work of fiction focused on Phyllis&#39;s adventures in atomic wonderland.

Phyllis chose well. But the crossroads metaphor was never really about choice. No-one was expected to pursue the path to nuclear annihilation. The crossroads demanded obedience to a specific vision of the future. It&#39;s this way ...or else.

&#39;Mapping Our Anzacs&#39; was created at a time of optimism, when it was thought that web technologies would open up the workings of government to new forms of public participation and transparency. But the dreams of &#39;government 2.0&#39; have faded, as information becomes ever more tightly controlled. In the GLAM sector, APIs have come and gone. Datasets created for long past hack events linger without updates, almost forgotten. New defensive measures aimed at taming the onslaught of AI scraper bots have imposed extra limits on access. Meanwhile a handful of tech oligarchs tell us what our future will be. This is the reality of progress.

The work of GLAM Labs, of GLAM innovation, has always been focused on expanding the realm of the possible – encouraging people to see differently, to think differently. This work struggles constantly with the many meanings of &#39;access&#39; – what use is data without good documentation, without permissive licences, without tools for analysis, without the skills of confidence to use those tools. It was this sort of struggle that motivated the [GLAM Workbench](https://glam-workbench.net/).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/557c57b82c.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide showing home page of the GLAM Workbench&#34;&gt;

I recently wrote a [potted introduction to the GLAM Workbench](https://updates.timsherratt.org/2025/06/05/glam-workbench-preprint-for-building.html) for a forthcoming publication on tool-making in the digital humanities. I won&#39;t read it all out, but I think it gives a pretty good overview of where things are.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-33.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Screenshot showing article in Zenodo&#34;&gt;

I suppose I want to emphasise though that the aim of the GLAM Workbench has always been to document possibilities – to expose researchers to the richness of GLAM data, to the new types of questions they can ask, and to the methods that are available to connect everything up.

In a world that erects multiple barriers of expertise, ownership, participation, and authority, there&#39;s power in simply knowing what&#39;s possible.

----

Back in 2010, at the first THATCamp Canberra, someone thanked me and said &#39;I&#39;ve found my people&#39;. That sense of belonging was always what made the National Digital Forum in New Zealand so special. I&#39;m not great at organisations – meetings make me anxious, and my email is a bin fire – but I do draw a lot of strength from the passions of like-minded people.

I don&#39;t really know what the future of the GLAM Workbench will be. I&#39;d like to be confident and optimistic, but the first half of last year &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;was pretty bleak&lt;/a&gt;, and left me wondering whether I should just walk away from all of it ­– go bushwalking, catch up on gardening, perhaps just be a historian again. I was saved by the GLAM Labs community. First of all, by the fabulous folk at the State Library of Victoria&#39;s LAB, particularly Paula Bray and Sotirios Alpanis. They gave me what I needed – fun data and wicked challenges. For a few months, I was back in my happy place, [creating new pathways through the SLV&#39;s place-based collections](https://slv.wraggelabs.com/).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-34.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide with screenshots showing the page that collects all the outputs from my SLV residency and an example of viewing photos by street in the CUA browser&#34;&gt;

The second boost was, of course, the invitation to be here today – to find out what&#39;s happening in labs around the world, to finally meet people I&#39;ve known online for years, to join in the excitement and, yes, to share the disappointments.

Perhaps the GLAM Workbench has done it&#39;s job. I think it&#39;s helped give people the confidence to dip a toe in the world of collections as data. It&#39;s also provided a useful model for GLAM organisations seeking to encourage new types of research. Perhaps its main value was always as an intervention – an invocation of possibilities that filled a particular gap at a particular moment in time. I&#39;ve always tried to wrap the GLAM Workbench in layers of documentation to enable any lasting value to be extracted as needed. Perhaps my focus should be to make sure that documentation is complete.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/glam-labs-2026-35.jpg&#34; width=&#34;600&#34; height=&#34;337&#34; alt=&#34;Slide with screenshot from the GLAM Workbench describing how organisations can create their own sections/repositories&#34;&gt;

But it&#39;s not a choice for me alone. The GLAM Workbench has [always welcomed contributions](https://glam-workbench.net/get-involved/developing-repositories/), so if you&#39;d like to carve out your own spaces, create your own sections, let me know! Perhaps the future of the GLAM Workbench will be shaped by the hands of others.

There&#39;s also a lot of cool GLAM data out there to play around with, and it really doesn&#39;t take much to get me excited about it. So perhaps I&#39;ll just continue to follow my enthusiasms and see where that takes things.

All of these possible futures are good. The choices aren&#39;t fixed -- they leave the conversation with history open and constructive. I&#39;m happy with that.

So greetings, thanks, and solidarity to all GLAM Labbers, past, present, and future. Despite all the setbacks and frustrations, your work matters. Take time to remember, to enjoy, and to celebrate.
</source:markdown>
    </item>
    
    <item>
      <title>The battle of the bots</title>
      <link>https://updates.timsherratt.org/2026/05/12/the-battle-of-the-bots.html</link>
      <pubDate>Tue, 12 May 2026 17:16:04 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/05/12/the-battle-of-the-bots.html</guid>
      <description>&lt;p&gt;I&amp;rsquo;ve spent a lot of time recently trying to protect my tools and resources from the relentless onslaught of AI scraper bots. I&amp;rsquo;ve documented much of the process here in case it&amp;rsquo;s of use to others (and to remind myself when I inevitably forget what I&amp;rsquo;ve done).&lt;/p&gt;
&lt;p&gt;Over the years, I&amp;rsquo;ve often made use of cloud services such as Heroku and Google Cloudrun to share my odd assortment of tools and experiments. In particular, Cloudrun has offered an easy way of publishing interactive  databases created using &lt;a href=&#34;https://datasette.io&#34;&gt;Datasette&lt;/a&gt; – one command and they&amp;rsquo;re online and open for exploration. This has worked (with a little fiddling) for databases of all sizes. The &lt;a href=&#34;https://glam-databases.net/gni/&#34;&gt;GLAM Name Index&lt;/a&gt;, for example, combines 10 databases totalling about 4 gigabytes of data. My only concern has been the amount of memory that Cloudrun demands to run Datasette databases. Throwing memory at things costs money, but because Cloudrun services can be configured to shut down when not in use, the overall hosting costs have remained manageable. Until now.&lt;/p&gt;
&lt;p&gt;Earlier this year, I noticed that the &lt;a href=&#34;https://resources.chineseaustralia.org/tung_wah_newspaper_index&#34;&gt;Tung Wah Newspaper Index&lt;/a&gt;, that ran using Datasette on Heroku, was having trouble – it kept falling over because it exceeded Heroku&amp;rsquo;s CPU and memory limits. The cause was a dramatic increase in traffic from AI scraper bots. I&amp;rsquo;ve written my share of data scrapers over the years, but I&amp;rsquo;ve always tried to avoid placing any undue burden on systems. The AI scraper bots are different, they&amp;rsquo;re indiscriminate and unyielding – they want everything, they want it now, and they don&amp;rsquo;t care how much damage they do. In the case of the Tung Wah Newspaper Index, the problem was both the overall number of requests, and the number of requests for computationally-intensive queries, such as facets. The only good thing was that the Heroku instance had a fixed cost, so the influx was killing the resource, but not costing me extra.&lt;/p&gt;
&lt;p&gt;I got the Tung Wah Newspaper Index running again by moving it to the Ubuntu server I run on Digital Ocean to host &lt;a href=&#34;https://wraggelabs.com&#34;&gt;wraggelabs.com&lt;/a&gt;. I fiddled around with my nginx settings to add a block list of known bots, throttled the number of requests per second, and switched off facets in Datasette. I also analysed the nginx logs to identify the IP address ranges producing the most traffic and blocked them as well. I knew that this might affect legitimate users, but I just wanted to get it working again for most people.&lt;/p&gt;
&lt;p&gt;Up until recently, the Datasette instances on Cloudrun didn&amp;rsquo;t seem to be attracting the same amount of traffic. But then I started getting alerts about unexpected billing increases. As I said above, Cloudrun instances can be configured to scale to zero when no requests are incoming – they switch themselves off when no one is using them. This conserves resources, and keeps costs down. But once the bots found my Datasette databases, the flood of requests was almost constant, so the services remained &amp;lsquo;on&amp;rsquo; and my bills went up. All together, my Datasette instances used to cost me around $30 a month, this skyrocketed overnight and would&amp;rsquo;ve cost me several hundred per month if I left things as they were.&lt;/p&gt;
&lt;p&gt;I was already using the &lt;a href=&#34;https://datasette.io/plugins/datasette-block-robots&#34;&gt;datasette-block-robots&lt;/a&gt; plugin to generate a &lt;code&gt;robots.txt&lt;/code&gt; file that told bots to go away. But the AI scraper bots just ignore these sorts of things. More drastic action was needed.&lt;/p&gt;
&lt;p&gt;Sysadmin stuff makes me nervous, and I don&amp;rsquo;t like to fiddle with things that seem to be working. I was reluctant to move my databases from Cloudrun because it was so convenient, and because I was worried that other platforms would struggle with the amount of data I was sharing. As a result, I spent a number of days investigating Google&amp;rsquo;s own bot protection services – this seemed to involve putting a load balancer in front of my instances, then attaching their &amp;lsquo;Cloud Armor&amp;rsquo; service to the load balancer. Of course, the more you paid, the more protection you got. I made a few attempts to get this going, but then gave up. The configuration itself was complex and confusing, but the main problem was that I felt I was just giving over more and more power to services that I didn&amp;rsquo;t really control or understand. The whole thing made me feel uneasy.&lt;/p&gt;
&lt;p&gt;I did a bit of local testing and realised that actual memory requirements for my Datasette databases was much less than Cloudrun demanded. This meant I could conceivably get them all up and running on a smallish virtual server, such as I&amp;rsquo;m using for wraggelabs.com. But what about the bots? I decided to give &lt;a href=&#34;https://anubis.techaro.lol&#34;&gt;Anubis&lt;/a&gt; a try. Some AI scraper bots identify themselves in the &lt;code&gt;User-Agent&lt;/code&gt; header field of their HTTP requests – they tell you their names. This makes them relatively easy to block. But other bots masquerade as normal web browsers, so their requests seem to come from humans. Anubis sniffs out disguised bots by asking them to complete a challenge that&amp;rsquo;s only possible in a real web browser. It seemed like a much more complete and consistent approach than my ad hoc efforts to block IP ranges.&lt;/p&gt;
&lt;p&gt;So the plan was:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;create a new domain for my databases&lt;/li&gt;
&lt;li&gt;create a new virtual server in a Digital Ocean droplet and attach the domain&lt;/li&gt;
&lt;li&gt;set up nginx and Anubis to manage web traffic and repel bot attacks&lt;/li&gt;
&lt;li&gt;migrate the databases, checking that they could comfortably fit within the server&amp;rsquo;s limits&lt;/li&gt;
&lt;li&gt;redirect the Cloudrun services to the new server&lt;/li&gt;
&lt;li&gt;monitor the bot traffic to make sure nothing was going to explode&lt;/li&gt;
&lt;/ul&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-05-12-16-47-42.png&#34; width=&#34;600&#34; height=&#34;349&#34; alt=&#34;Screenshot of the glam-databases.net homepage&#34;&gt;&lt;figcaption&gt;&lt;a ref=&#34;https://glam-databases.net&#34;&gt;A new home for my GLAM databases&lt;/a&gt;, with built-in bot protection&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;It&amp;rsquo;s all now done, the databases seem happy in their new home, Anubis is busy blocking bots, and my costs are contained. You can visit the databases at &lt;a href=&#34;https://glam-databases.net&#34;&gt;glam-databases.net&lt;/a&gt;. The gory details are below&amp;hellip;&lt;/p&gt;
&lt;h2 id=&#34;the-technical-details&#34;&gt;The technical details&lt;/h2&gt;
&lt;p&gt;I registered glam-databases.net and set up an Ubuntu server in a Sydney-located Digital Ocean droplet. I did the usual &lt;a href=&#34;https://www.digitalocean.com/community/tutorials/initial-server-setup-with-ubuntu&#34;&gt;security stuff&lt;/a&gt;, like setting up a firewall and a non-root user, and then &lt;a href=&#34;https://www.digitalocean.com/community/tutorials/how-to-install-nginx-on-ubuntu-22-04&#34;&gt;installed nginx and SSH&lt;/a&gt; using the Digital Ocean guides.&lt;/p&gt;
&lt;p&gt;I installed Anubis using &lt;a href=&#34;https://anubis.techaro.lol/docs/admin/native-install&#34;&gt;the instructions on the website&lt;/a&gt;. However, I wanted to use Unix sockets to connect everything up and ran into a few problems getting Anubis to write &lt;code&gt;.sock&lt;/code&gt; files. &lt;a href=&#34;https://github.com/TecharoHQ/anubis/issues/1583&#34;&gt;This issue&lt;/a&gt; helped me understand that I needed to create a directory with the name of my Anubis service in &lt;code&gt;/run/anubis&lt;/code&gt;. I eventually figured out that I also needed to add &lt;code&gt;SOCKET_MODE=0777&lt;/code&gt; to the Anubis config file to get the permissions right.&lt;/p&gt;
&lt;p&gt;The other thing I struggled with a little was setting up logging. Once again it was mostly a permissions thing. &lt;a href=&#34;https://anubis.techaro.lol/docs/admin/policies#file-sink&#34;&gt;The Anubis documentation&lt;/a&gt; explains that if you&amp;rsquo;re running Anubis via &lt;code&gt;systemd&lt;/code&gt; you need to add a drop-in unit to allow Anubis to write to &lt;code&gt;/var/log&lt;/code&gt;, however, I found that the drop-in needed to be saved to &lt;code&gt;/etc/systemd/system/anubis@instance-name.service.d/&lt;/code&gt; rather than &lt;code&gt;/etc/systemd/anubis@instance-name.service.d/&lt;/code&gt; as stated in the docs (notice the extra &lt;code&gt;system&lt;/code&gt; directory).&lt;/p&gt;
&lt;p&gt;Once the socket stuff was sorted, getting Anubis to talk to nginx was pretty straightforward, and I mostly just followed &lt;a href=&#34;https://anubis.techaro.lol/docs/admin/environments/nginx&#34;&gt;the Anubis documentation&lt;/a&gt; to make the necessary changes to the nginx config files. I&amp;rsquo;d set up a basic, static homepage for the site, so I tried accessing that to check that Anubis was in fact working. The Anubis challenge page was displayed as expected, and I could see the challenge details in the logs. Yay!&lt;/p&gt;
&lt;p&gt;Next it was time to start moving my databases across. The first step was to ensure that each of my Datasette repositories included a &lt;code&gt;requirements.in&lt;/code&gt; file, listing the necessary Python packages – this varied a bit across projects depending on the Datasette plugins they used, but the minimum was &lt;code&gt;datasette&lt;/code&gt; and the &lt;code&gt;datasette-block-robots&lt;/code&gt; plugin. To improve performance I &lt;a href=&#34;https://docs.datasette.io/en/stable/performance.html#using-datasette-inspect&#34;&gt;used &lt;code&gt;datasette inspect&lt;/code&gt;&lt;/a&gt; to calculate table row counts and write them to a json file in the repository. The counts file is loaded at run time using the &lt;code&gt;--inspect-file&lt;/code&gt; parameter.&lt;/p&gt;
&lt;p&gt;I created a directory for each of the databases on the new server, and used &lt;code&gt;rsync&lt;/code&gt; to copy the database repositories across. Then in each of the database directories I created a virtual environment using pyenv, and installed all the necessary Python packages using pip-tools and the &lt;code&gt;requirements.in&lt;/code&gt; file.&lt;/p&gt;
&lt;p&gt;To get the databases running, I mostly just followed the Datasette documentation which provides information on &lt;a href=&#34;https://docs.datasette.io/en/stable/deploying.html#running-datasette-using-systemd&#34;&gt;running Datasette using systemd&lt;/a&gt; and &lt;a href=&#34;https://docs.datasette.io/en/stable/deploying.html#nginx-proxy-configuration&#34;&gt;configuring nginx&lt;/a&gt;. I used the &lt;code&gt;--uds&lt;/code&gt; and &lt;code&gt;--base_url&lt;/code&gt; settings to connect via unix sockets and deploy each database at a different url prefix. The only problem I found is that the &lt;code&gt;base_url&lt;/code&gt; setting can be a little buggy – for example, links to json versions of pages sometimes included the base path prefix twice. Rather than fiddle about in the Datasette code, I decided to &amp;lsquo;fix&amp;rsquo; this using &lt;code&gt;rewrite&lt;/code&gt; directives in the nginx configuration file, for example: &lt;code&gt;rewrite ^/gni/gni/(.+)$ /gni/$1 break;&lt;/code&gt; I&amp;rsquo;m sure there are better solutions, but I figured this would do the job until things got fixed upstream.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s an example of the systemctl &lt;code&gt;service&lt;/code&gt; files I created for each database (with the key redacted):&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;[Unit]
Description=Sand and MacDougall&#39;s Directories (Victoria) Datasette
After=network.target

[Service]
User=tim
Group=www-data
Environment=DATASETTE_SECRET=**********************************
WorkingDirectory=/home/tim/databases/victoria-sands-mac-dirs
ExecStart=/home/tim/.pyenv/versions/victoria-sands-mac-dirs/bin/datasette serve -i sands-mcdougalls-directories-victoria.db --uds /tmp/victoria-sands-mac-dirs.sock   --inspect-file=counts.json --template-dir templates --static static:static  -m metadata.json --setting base_url /vic-smd/  --setting facet_time_limit_ms 10000 --setting suggest_facets off --setting allow_csv_stream off
Restart=on-failure

[Install]
WantedBy=multi-user.target
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The parameters used when starting Datasette vary a bit depending on the project (for example some don&amp;rsquo;t have custom templates) but in general they are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;-i&lt;/code&gt; run in immutable mode for better performance&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--uds&lt;/code&gt; path to unix socket&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--inspect-file&lt;/code&gt; path to precalculated row counts&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--template-dir&lt;/code&gt; custom template directory&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--static&lt;/code&gt; directory and url prefix for static files&lt;/li&gt;
&lt;li&gt;&lt;code&gt;-m&lt;/code&gt; path to metadata file&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--setting base_url&lt;/code&gt; url prefix pointing to this database&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--setting facet_time_limit_ms&lt;/code&gt; maximum time allowed for facet requests&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--setting suggest_facets off&lt;/code&gt; don&amp;rsquo;t provide example facets (both for performance and to avoid providing new paths for bots to find)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;--setting allow_csv_stream off&lt;/code&gt; for performance&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I created a service file like this for every database and saved it to the &lt;code&gt;/etc/systemd/system/&lt;/code&gt; directory.&lt;/p&gt;
&lt;p&gt;For each database I created two new sections in the nginx config file. First in the main &lt;code&gt;server&lt;/code&gt; block I added a &lt;code&gt;location&lt;/code&gt; entry that directed requests to the database&amp;rsquo;s url prefix. For example:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;location /vic-smd {
        rewrite ^/vic-smd/vic-smd/sands-mcdougalls-directories-victoria/(.+)$ /vic-smd/sands-mcdougalls-directories-victoria/$1 break;
        proxy_pass http://victoria-sands-mac-dirs/vic-smd;
        proxy_set_header Host $host;
}
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Then I created an &lt;code&gt;upstream&lt;/code&gt; entry for each service that send these requests to the database socket.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;upstream victoria-sands-mac-dirs {
	server unix:/tmp/victoria-sands-mac-dirs.sock;
}
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Of course, after all the changes I ran &lt;code&gt;sudo systemctl start&lt;/code&gt; or &lt;code&gt;restart&lt;/code&gt;  and checked that things were ok using &lt;code&gt;sudo systemctl status&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Once I had migrated all the databases, I left things alone for a day to make sure they were comfortably settled into their new home and that the server didn&amp;rsquo;t explode. I monitored the system stats, and found that the memory consumption was much lower than it had been on Cloudrun. So all seemed good.&lt;/p&gt;
&lt;p&gt;Then it was time to redirect the Cloudrun urls to the new database locations. When I created the Cloudrun instances, there was no option to add custom domains to instances in Australia. In most cases, I created redirects from the GLAM Workbench that I encouraged people to use and share, but I noticed that people were often saving the ugly Cloudrun urls instead. This meant that I couldn&amp;rsquo;t just switch the domain config, I had to replace the Cloudrun instances themselves with redirects. This ended up being much easier than I expected. Using &lt;a href=&#34;https://github.com/MorbZ/docker-web-redirect&#34;&gt;docker-web-direct&lt;/a&gt; I could spin up little nginx proxies to replace the Cloudrun databases. For example to redirect the Tasmanian Post Office Directories, I used the Cloudrun CLI to send the following command:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;gcloud run deploy tasmanian-post-office-directories --image morbz/docker-web-redirect --allow-unauthenticated --set-env-vars &amp;quot;REDIRECT_TARGET=https://glam-databases.net/tas-pod&amp;quot;
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The key things are the service name, in this case &lt;code&gt;tasmanian-post-office-directories&lt;/code&gt;, and the new url &lt;code&gt;https://glam-databases.net/tas-pod&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;I repeated this for each of the databases on Cloudrun. Of course, redirecting the urls to help people find the new locations also opened the door to the scraper bots, and they quickly muscled their way through. I was monitoring both the Anubis log file and the system stats as I redirected each database – the traffic went up, but everything was pretty stable, and Anubis seemed to be working.&lt;/p&gt;
&lt;figure&gt;&lt;video src=&#34;https://cdn.uploads.micro.mov/8371/2026/bots-2026-05-08-10.54.54/playlist.m3u8&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34;&gt;&lt;/video&gt;&lt;figcaption&gt;A realtime snapshot of the Anubis log file&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Checking the nginx log I noticed that one bot, &lt;code&gt;meta-webindexer&lt;/code&gt;, was hitting one of the databases hard and wasn&amp;rsquo;t being blocked by Anubis. There&amp;rsquo;s an &lt;a href=&#34;https://github.com/TecharoHQ/anubis/issues/1565&#34;&gt;open issue&lt;/a&gt; in the Anubis GitHub repo to add this bot to the blocklist, but it was easy to modify the default bot policy file to repel it:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;  - name: meta-webindexer
    user_agent_regex: meta-webindexer
    action: DENY
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;Datasette has a built-in JSON API and I use this to pull data into things like my SLV apps. For this to continue working I had to open a pathway in the Anubis bot policy file:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;  - name: allow-api-requests
    action: ALLOW
    expression:
      all:
        - &#39;&amp;quot;Accept&amp;quot; in headers&#39;
        - &#39;headers[&amp;quot;Accept&amp;quot;].contains(&amp;quot;application/json&amp;quot;)&#39;
        - &#39;path.contains(&amp;quot;.json&amp;quot;)
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;This means that requests with &lt;code&gt;.json&lt;/code&gt; in the url and &lt;code&gt;application/json&lt;/code&gt; in the accept headers will not be challenged by Anubis. I then modified some of the apps to make sure they actually were sending the correct headers, and it worked!&lt;/p&gt;
&lt;h2 id=&#34;next-steps&#34;&gt;Next steps&lt;/h2&gt;
&lt;p&gt;Things have gone pretty smoothly so far, but I&amp;rsquo;ll be keeping my on the stats. Now that I understand how Anubis works, I also want to install it on Wragge Labs and remove the ad hoc blocks I put in place earlier.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ll leave the redirects running on Cloudrun for a while, but not forever. If you use any of these databases, please update your bookmarks and links!&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://glam-databases.net/nsw-pod/&#34;&gt;NSW Post Office Directories&lt;/a&gt; (54 volumes from 1886 to 1950)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://glam-databases.net/syd-td/&#34;&gt;Sydney Telephone Directories&lt;/a&gt; (44 volumes from 1926 to 1954)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://glam-databases.net/tas-pod/&#34;&gt;Tasmanian Post Office Directories&lt;/a&gt; (54 volumes from 1890 to 1948)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://glam-databases.net/vic-smd/&#34;&gt;Sands &amp;amp; McDougall&amp;rsquo;s Directories, Victoria&lt;/a&gt; (24 volumes from 1860 to 1974)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://glam-databases.net/gni/&#34;&gt;GLAM Name Index Search&lt;/a&gt; (12 million records from 293 datasets)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;I&amp;rsquo;m relieved to have brought my Cloudrun costs under control, but all of this still costs money. Over recent years, I&amp;rsquo;ve been lucky to have the support of my GitHub sponsors to help me pay my cloud hosting bills. I&amp;rsquo;m planning on moving away from GitHub eventually and I&amp;rsquo;ve set up a &lt;a href=&#34;https://liberapay.com/wragge/&#34;&gt;brand new profile on LiberaPay&lt;/a&gt;. They&amp;rsquo;re an open source, community managed organisation that doesn&amp;rsquo;t take a cut of your donations, so if you want to help me keep things online you can set up a sponsorship there.&lt;/p&gt;
</description>
      <source:markdown>I&#39;ve spent a lot of time recently trying to protect my tools and resources from the relentless onslaught of AI scraper bots. I&#39;ve documented much of the process here in case it&#39;s of use to others (and to remind myself when I inevitably forget what I&#39;ve done).

Over the years, I&#39;ve often made use of cloud services such as Heroku and Google Cloudrun to share my odd assortment of tools and experiments. In particular, Cloudrun has offered an easy way of publishing interactive  databases created using [Datasette](https://datasette.io) – one command and they&#39;re online and open for exploration. This has worked (with a little fiddling) for databases of all sizes. The [GLAM Name Index](https://glam-databases.net/gni/), for example, combines 10 databases totalling about 4 gigabytes of data. My only concern has been the amount of memory that Cloudrun demands to run Datasette databases. Throwing memory at things costs money, but because Cloudrun services can be configured to shut down when not in use, the overall hosting costs have remained manageable. Until now.

Earlier this year, I noticed that the [Tung Wah Newspaper Index](https://resources.chineseaustralia.org/tung_wah_newspaper_index), that ran using Datasette on Heroku, was having trouble – it kept falling over because it exceeded Heroku&#39;s CPU and memory limits. The cause was a dramatic increase in traffic from AI scraper bots. I&#39;ve written my share of data scrapers over the years, but I&#39;ve always tried to avoid placing any undue burden on systems. The AI scraper bots are different, they&#39;re indiscriminate and unyielding – they want everything, they want it now, and they don&#39;t care how much damage they do. In the case of the Tung Wah Newspaper Index, the problem was both the overall number of requests, and the number of requests for computationally-intensive queries, such as facets. The only good thing was that the Heroku instance had a fixed cost, so the influx was killing the resource, but not costing me extra.

I got the Tung Wah Newspaper Index running again by moving it to the Ubuntu server I run on Digital Ocean to host [wraggelabs.com](https://wraggelabs.com). I fiddled around with my nginx settings to add a block list of known bots, throttled the number of requests per second, and switched off facets in Datasette. I also analysed the nginx logs to identify the IP address ranges producing the most traffic and blocked them as well. I knew that this might affect legitimate users, but I just wanted to get it working again for most people.

Up until recently, the Datasette instances on Cloudrun didn&#39;t seem to be attracting the same amount of traffic. But then I started getting alerts about unexpected billing increases. As I said above, Cloudrun instances can be configured to scale to zero when no requests are incoming – they switch themselves off when no one is using them. This conserves resources, and keeps costs down. But once the bots found my Datasette databases, the flood of requests was almost constant, so the services remained &#39;on&#39; and my bills went up. All together, my Datasette instances used to cost me around $30 a month, this skyrocketed overnight and would&#39;ve cost me several hundred per month if I left things as they were.

I was already using the [datasette-block-robots](https://datasette.io/plugins/datasette-block-robots) plugin to generate a `robots.txt` file that told bots to go away. But the AI scraper bots just ignore these sorts of things. More drastic action was needed.

Sysadmin stuff makes me nervous, and I don&#39;t like to fiddle with things that seem to be working. I was reluctant to move my databases from Cloudrun because it was so convenient, and because I was worried that other platforms would struggle with the amount of data I was sharing. As a result, I spent a number of days investigating Google&#39;s own bot protection services – this seemed to involve putting a load balancer in front of my instances, then attaching their &#39;Cloud Armor&#39; service to the load balancer. Of course, the more you paid, the more protection you got. I made a few attempts to get this going, but then gave up. The configuration itself was complex and confusing, but the main problem was that I felt I was just giving over more and more power to services that I didn&#39;t really control or understand. The whole thing made me feel uneasy.

I did a bit of local testing and realised that actual memory requirements for my Datasette databases was much less than Cloudrun demanded. This meant I could conceivably get them all up and running on a smallish virtual server, such as I&#39;m using for wraggelabs.com. But what about the bots? I decided to give [Anubis](https://anubis.techaro.lol) a try. Some AI scraper bots identify themselves in the `User-Agent` header field of their HTTP requests – they tell you their names. This makes them relatively easy to block. But other bots masquerade as normal web browsers, so their requests seem to come from humans. Anubis sniffs out disguised bots by asking them to complete a challenge that&#39;s only possible in a real web browser. It seemed like a much more complete and consistent approach than my ad hoc efforts to block IP ranges.

So the plan was:

- create a new domain for my databases
- create a new virtual server in a Digital Ocean droplet and attach the domain
- set up nginx and Anubis to manage web traffic and repel bot attacks
- migrate the databases, checking that they could comfortably fit within the server&#39;s limits
- redirect the Cloudrun services to the new server
- monitor the bot traffic to make sure nothing was going to explode

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-05-12-16-47-42.png&#34; width=&#34;600&#34; height=&#34;349&#34; alt=&#34;Screenshot of the glam-databases.net homepage&#34;&gt;&lt;figcaption&gt;&lt;a ref=&#34;https://glam-databases.net&#34;&gt;A new home for my GLAM databases&lt;/a&gt;, with built-in bot protection&lt;/figcaption&gt;&lt;/figure&gt;

It&#39;s all now done, the databases seem happy in their new home, Anubis is busy blocking bots, and my costs are contained. You can visit the databases at [glam-databases.net](https://glam-databases.net). The gory details are below...

## The technical details

I registered glam-databases.net and set up an Ubuntu server in a Sydney-located Digital Ocean droplet. I did the usual [security stuff](https://www.digitalocean.com/community/tutorials/initial-server-setup-with-ubuntu), like setting up a firewall and a non-root user, and then [installed nginx and SSH](https://www.digitalocean.com/community/tutorials/how-to-install-nginx-on-ubuntu-22-04) using the Digital Ocean guides.

I installed Anubis using [the instructions on the website](https://anubis.techaro.lol/docs/admin/native-install). However, I wanted to use Unix sockets to connect everything up and ran into a few problems getting Anubis to write `.sock` files. [This issue](https://github.com/TecharoHQ/anubis/issues/1583) helped me understand that I needed to create a directory with the name of my Anubis service in `/run/anubis`. I eventually figured out that I also needed to add `SOCKET_MODE=0777` to the Anubis config file to get the permissions right.

The other thing I struggled with a little was setting up logging. Once again it was mostly a permissions thing. [The Anubis documentation](https://anubis.techaro.lol/docs/admin/policies#file-sink) explains that if you&#39;re running Anubis via `systemd` you need to add a drop-in unit to allow Anubis to write to `/var/log`, however, I found that the drop-in needed to be saved to `/etc/systemd/system/anubis@instance-name.service.d/` rather than `/etc/systemd/anubis@instance-name.service.d/` as stated in the docs (notice the extra `system` directory).

Once the socket stuff was sorted, getting Anubis to talk to nginx was pretty straightforward, and I mostly just followed [the Anubis documentation](https://anubis.techaro.lol/docs/admin/environments/nginx) to make the necessary changes to the nginx config files. I&#39;d set up a basic, static homepage for the site, so I tried accessing that to check that Anubis was in fact working. The Anubis challenge page was displayed as expected, and I could see the challenge details in the logs. Yay! 

Next it was time to start moving my databases across. The first step was to ensure that each of my Datasette repositories included a `requirements.in` file, listing the necessary Python packages – this varied a bit across projects depending on the Datasette plugins they used, but the minimum was `datasette` and the `datasette-block-robots` plugin. To improve performance I [used `datasette inspect`](https://docs.datasette.io/en/stable/performance.html#using-datasette-inspect) to calculate table row counts and write them to a json file in the repository. The counts file is loaded at run time using the `--inspect-file` parameter.

I created a directory for each of the databases on the new server, and used `rsync` to copy the database repositories across. Then in each of the database directories I created a virtual environment using pyenv, and installed all the necessary Python packages using pip-tools and the `requirements.in` file.

To get the databases running, I mostly just followed the Datasette documentation which provides information on [running Datasette using systemd](https://docs.datasette.io/en/stable/deploying.html#running-datasette-using-systemd) and [configuring nginx](https://docs.datasette.io/en/stable/deploying.html#nginx-proxy-configuration). I used the `--uds` and `--base_url` settings to connect via unix sockets and deploy each database at a different url prefix. The only problem I found is that the `base_url` setting can be a little buggy – for example, links to json versions of pages sometimes included the base path prefix twice. Rather than fiddle about in the Datasette code, I decided to &#39;fix&#39; this using `rewrite` directives in the nginx configuration file, for example: `rewrite ^/gni/gni/(.+)$ /gni/$1 break;` I&#39;m sure there are better solutions, but I figured this would do the job until things got fixed upstream.

Here&#39;s an example of the systemctl `service` files I created for each database (with the key redacted):

```
[Unit]
Description=Sand and MacDougall&#39;s Directories (Victoria) Datasette
After=network.target

[Service]
User=tim
Group=www-data
Environment=DATASETTE_SECRET=**********************************
WorkingDirectory=/home/tim/databases/victoria-sands-mac-dirs
ExecStart=/home/tim/.pyenv/versions/victoria-sands-mac-dirs/bin/datasette serve -i sands-mcdougalls-directories-victoria.db --uds /tmp/victoria-sands-mac-dirs.sock   --inspect-file=counts.json --template-dir templates --static static:static  -m metadata.json --setting base_url /vic-smd/  --setting facet_time_limit_ms 10000 --setting suggest_facets off --setting allow_csv_stream off
Restart=on-failure

[Install]
WantedBy=multi-user.target
```
The parameters used when starting Datasette vary a bit depending on the project (for example some don&#39;t have custom templates) but in general they are:

- `-i` run in immutable mode for better performance
-  `--uds` path to unix socket
- `--inspect-file` path to precalculated row counts
- `--template-dir` custom template directory
- `--static` directory and url prefix for static files
- `-m` path to metadata file
- `--setting base_url` url prefix pointing to this database
- `--setting facet_time_limit_ms` maximum time allowed for facet requests
- `--setting suggest_facets off` don&#39;t provide example facets (both for performance and to avoid providing new paths for bots to find)
- `--setting allow_csv_stream off` for performance

I created a service file like this for every database and saved it to the `/etc/systemd/system/` directory.

For each database I created two new sections in the nginx config file. First in the main `server` block I added a `location` entry that directed requests to the database&#39;s url prefix. For example:

```
location /vic-smd {
        rewrite ^/vic-smd/vic-smd/sands-mcdougalls-directories-victoria/(.+)$ /vic-smd/sands-mcdougalls-directories-victoria/$1 break;
        proxy_pass http://victoria-sands-mac-dirs/vic-smd;
        proxy_set_header Host $host;
}
```
Then I created an `upstream` entry for each service that send these requests to the database socket.

```
upstream victoria-sands-mac-dirs {
	server unix:/tmp/victoria-sands-mac-dirs.sock;
}
```
Of course, after all the changes I ran `sudo systemctl start` or `restart`  and checked that things were ok using `sudo systemctl status`.

Once I had migrated all the databases, I left things alone for a day to make sure they were comfortably settled into their new home and that the server didn&#39;t explode. I monitored the system stats, and found that the memory consumption was much lower than it had been on Cloudrun. So all seemed good.

Then it was time to redirect the Cloudrun urls to the new database locations. When I created the Cloudrun instances, there was no option to add custom domains to instances in Australia. In most cases, I created redirects from the GLAM Workbench that I encouraged people to use and share, but I noticed that people were often saving the ugly Cloudrun urls instead. This meant that I couldn&#39;t just switch the domain config, I had to replace the Cloudrun instances themselves with redirects. This ended up being much easier than I expected. Using [docker-web-direct](https://github.com/MorbZ/docker-web-redirect) I could spin up little nginx proxies to replace the Cloudrun databases. For example to redirect the Tasmanian Post Office Directories, I used the Cloudrun CLI to send the following command:

```
gcloud run deploy tasmanian-post-office-directories --image morbz/docker-web-redirect --allow-unauthenticated --set-env-vars &#34;REDIRECT_TARGET=https://glam-databases.net/tas-pod&#34;
```
The key things are the service name, in this case `tasmanian-post-office-directories`, and the new url `https://glam-databases.net/tas-pod`. 

I repeated this for each of the databases on Cloudrun. Of course, redirecting the urls to help people find the new locations also opened the door to the scraper bots, and they quickly muscled their way through. I was monitoring both the Anubis log file and the system stats as I redirected each database – the traffic went up, but everything was pretty stable, and Anubis seemed to be working.

&lt;figure&gt;&lt;video src=&#34;https://cdn.uploads.micro.mov/8371/2026/bots-2026-05-08-10.54.54/playlist.m3u8&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34;&gt;&lt;/video&gt;&lt;figcaption&gt;A realtime snapshot of the Anubis log file&lt;/figcaption&gt;&lt;/figure&gt;

Checking the nginx log I noticed that one bot, `meta-webindexer`, was hitting one of the databases hard and wasn&#39;t being blocked by Anubis. There&#39;s an [open issue](https://github.com/TecharoHQ/anubis/issues/1565) in the Anubis GitHub repo to add this bot to the blocklist, but it was easy to modify the default bot policy file to repel it:

```
  - name: meta-webindexer
    user_agent_regex: meta-webindexer
    action: DENY
```
Datasette has a built-in JSON API and I use this to pull data into things like my SLV apps. For this to continue working I had to open a pathway in the Anubis bot policy file:

```
  - name: allow-api-requests
    action: ALLOW
    expression:
      all:
        - &#39;&#34;Accept&#34; in headers&#39;
        - &#39;headers[&#34;Accept&#34;].contains(&#34;application/json&#34;)&#39;
        - &#39;path.contains(&#34;.json&#34;)
```
This means that requests with `.json` in the url and `application/json` in the accept headers will not be challenged by Anubis. I then modified some of the apps to make sure they actually were sending the correct headers, and it worked!

## Next steps

Things have gone pretty smoothly so far, but I&#39;ll be keeping my on the stats. Now that I understand how Anubis works, I also want to install it on Wragge Labs and remove the ad hoc blocks I put in place earlier.

I&#39;ll leave the redirects running on Cloudrun for a while, but not forever. If you use any of these databases, please update your bookmarks and links!

- [NSW Post Office Directories](https://glam-databases.net/nsw-pod/) (54 volumes from 1886 to 1950)
- [Sydney Telephone Directories](https://glam-databases.net/syd-td/) (44 volumes from 1926 to 1954)
- [Tasmanian Post Office Directories](https://glam-databases.net/tas-pod/) (54 volumes from 1890 to 1948)
- [Sands &amp; McDougall&#39;s Directories, Victoria](https://glam-databases.net/vic-smd/) (24 volumes from 1860 to 1974)
- [GLAM Name Index Search](https://glam-databases.net/gni/) (12 million records from 293 datasets)

I&#39;m relieved to have brought my Cloudrun costs under control, but all of this still costs money. Over recent years, I&#39;ve been lucky to have the support of my GitHub sponsors to help me pay my cloud hosting bills. I&#39;m planning on moving away from GitHub eventually and I&#39;ve set up a [brand new profile on LiberaPay](https://liberapay.com/wragge/). They&#39;re an open source, community managed organisation that doesn&#39;t take a cut of your donations, so if you want to help me keep things online you can set up a sponsorship there.






</source:markdown>
    </item>
    
    <item>
      <title>Generosity in practice – a chat with Paula Bray at the State Library of Victoria</title>
      <link>https://updates.timsherratt.org/2026/03/16/generosity-in-practice-a-chat.html</link>
      <pubDate>Mon, 16 Mar 2026 14:14:17 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/03/16/generosity-in-practice-a-chat.html</guid>
      <description>&lt;p&gt;While I was in Melbourne during my time as &lt;a href=&#34;https://lab.slv.vic.gov.au/team/tim-sherratt&#34;&gt;Creative Technologist-in-Residence at the State Library of Victorian LAB&lt;/a&gt;, I had a conversation with Paula Bray for the LAB&amp;rsquo;s podcast series. Paula is the SLV&amp;rsquo;s Chief Digital Officer, and has long championed the importance of digital innovation in the GLAM sector. It was fun to chat about stuff that I&amp;rsquo;ve been doing for the last 30 years, any why openness and generosity is important in working with GLAM collections. You can &lt;a href=&#34;https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt/interview&#34;&gt;listen to our conversation on the LAB site&lt;/a&gt;.&lt;/p&gt;
</description>
      <source:markdown>While I was in Melbourne during my time as [Creative Technologist-in-Residence at the State Library of Victorian LAB](https://lab.slv.vic.gov.au/team/tim-sherratt), I had a conversation with Paula Bray for the LAB&#39;s podcast series. Paula is the SLV&#39;s Chief Digital Officer, and has long championed the importance of digital innovation in the GLAM sector. It was fun to chat about stuff that I&#39;ve been doing for the last 30 years, any why openness and generosity is important in working with GLAM collections. You can [listen to our conversation on the LAB site](https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt/interview).


</source:markdown>
    </item>
    
    <item>
      <title>Zotero translator for Libraries Tasmania updated!</title>
      <link>https://updates.timsherratt.org/2026/03/10/zotero-translator-for-libraries-tasmania.html</link>
      <pubDate>Tue, 10 Mar 2026 09:31:36 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/03/10/zotero-translator-for-libraries-tasmania.html</guid>
      <description>&lt;p&gt;The &lt;a href=&#34;https://www.zotero.org&#34;&gt;Zotero&lt;/a&gt; translator for &lt;a href=&#34;https://libraries.tas.gov.au&#34;&gt;Libraries Tasmania&lt;/a&gt; has been updated, fixing a problem with attaching images of digitised resources. The fix is in the main Zotero repository now, so it should find its way to your computer automatically.&lt;/p&gt;
&lt;p&gt;I created the first version of the Libraries Tasmania translator back in 2022 – &lt;a href=&#34;https://updates.timsherratt.org/2022/07/14/calling-all-tasmanian.html&#34;&gt;this post describes what it does&lt;/a&gt;. It works across all three sections of the catalogue, including the archives, and names index. The translator captures metadata, PDFs, and images from records, including things like digitised pages from convict records. This makes it easy for researchers to assemble their own datasets of Tasmanian records in Zotero, where they can add notes and annotations, or share with colleagues.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/zotero-librariestas.png&#34; width=&#34;600&#34; height=&#34;382&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Capture images and metadata from the Libraries Tasmania catalogue using Zotero&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The update was necessary because Libraries Tasmania changed the way some digitised resources were displayed and downloaded. Keeping Zotero translators working across system updates can take a bit of work! I also took the opportunity to update the code to meet current Zotero guidelines and clean up a few lingering problems. If you notice any oddities, please let me know.&lt;/p&gt;
&lt;p&gt;There are now at least &lt;a href=&#34;https://updates.timsherratt.org/2024/08/22/new-zotero-translators.html#zotero-and-australian-glams&#34;&gt;8 custom translators&lt;/a&gt; to help you work with Australian GLAM collections.&lt;/p&gt;
</description>
      <source:markdown>The [Zotero](https://www.zotero.org) translator for [Libraries Tasmania](https://libraries.tas.gov.au) has been updated, fixing a problem with attaching images of digitised resources. The fix is in the main Zotero repository now, so it should find its way to your computer automatically.

I created the first version of the Libraries Tasmania translator back in 2022 – [this post describes what it does](https://updates.timsherratt.org/2022/07/14/calling-all-tasmanian.html). It works across all three sections of the catalogue, including the archives, and names index. The translator captures metadata, PDFs, and images from records, including things like digitised pages from convict records. This makes it easy for researchers to assemble their own datasets of Tasmanian records in Zotero, where they can add notes and annotations, or share with colleagues.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/zotero-librariestas.png&#34; width=&#34;600&#34; height=&#34;382&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Capture images and metadata from the Libraries Tasmania catalogue using Zotero&lt;/figcaption&gt;&lt;/figure&gt;

The update was necessary because Libraries Tasmania changed the way some digitised resources were displayed and downloaded. Keeping Zotero translators working across system updates can take a bit of work! I also took the opportunity to update the code to meet current Zotero guidelines and clean up a few lingering problems. If you notice any oddities, please let me know.

There are now at least [8 custom translators](https://updates.timsherratt.org/2024/08/22/new-zotero-translators.html#zotero-and-australian-glams) to help you work with Australian GLAM collections. 
</source:markdown>
    </item>
    
    <item>
      <title>Exploring georeferenced maps from the SLV collection</title>
      <link>https://updates.timsherratt.org/2026/02/12/exploring-georeferenced-maps-from-the.html</link>
      <pubDate>Thu, 12 Feb 2026 22:05:06 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/02/12/exploring-georeferenced-maps-from-the.html</guid>
      <description>&lt;p&gt;I&amp;rsquo;m in the process of tying up all the documentation relating to my time as &lt;a href=&#34;https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt/people-place-library-data-tim-sherratt&#34;&gt;Creative Technologist-in-Residence at the State Library of Victoria LAB&lt;/a&gt;. But as I was looking through &lt;a href=&#34;https://slv.wraggelabs.com&#34;&gt;the list of outputs&lt;/a&gt;, I realised I&amp;rsquo;d never written anything about the interface I created to explore georeferenced maps from the SLV collection.&lt;/p&gt;
&lt;p&gt;I also remembered that there were a few improvements I wanted to make to the interface. So instead of spending a few hours writing up a blog post, I&amp;rsquo;ve spent several days completely overhauling the &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Georeferenced Maps Explorer&lt;/a&gt;. I&amp;rsquo;m pretty happy with how it&amp;rsquo;s working now. &lt;strong&gt;&lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Have a play!&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/prom-maps.png&#34; width=&#34;600&#34; height=&#34;354&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Wilson&#39;s Prom made up of a patchwork of georeferenced maps and aerial photographs using the Georefrenced Maps Explorer. &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Try it now!&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;To get started, just click on the basemap. Details of all georeferenced maps within 50km of your selected point will be displayed in the right-hand column. As you move your mouse over the list of results, the boundaries of the georeferenced maps will be displayed on the basemap. This gives you a preview of their location and size. Click on one of the results to display the georeferenced map as a layer on top of the modern basemap.&lt;/p&gt;
&lt;figure&gt;&lt;video src=&#34;https://cdn.uploads.micro.mov/8371/2026/video-2026-02-12-13-39-23/playlist.m3u8&#34; poster=&#34;https://cdn.uploads.micro.blog/8371/2026/frames/1679959-0-ad6b9b.jpg&#34; width=&#34;1920&#34; height=&#34;1080&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34;&gt;&lt;/video&gt;&lt;figcaption&gt;Hover over a result to see the map boundaries&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;You can add as many maps as you like. If your selected maps overlap, you can change the order in which they&amp;rsquo;re shown. Click on the layers icon in the top left of the basemap. You&amp;rsquo;ll see a list of the maps that are currently displayed. Use the arrow buttons to move a map backwards or forwards. You can also use the sliders to adjust the opacity of each map. This can make it easier to examine the relationship between maps. For example, you might want to compare the features of a historic map with those of the underlying basemap.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/photomaps.png&#34; width=&#34;600&#34; height=&#34;363&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Stitch together multiple maps like this series of seven photomaps, and change the opacity to see the features underneath&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;The Explorer&amp;rsquo;s url updates with every selection you make, so you can bookmark or share a url to return to the same position and collection of maps. For example, &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/?lat=-38.96774552450964&amp;amp;lon=146.39004985426214&amp;amp;map_id=215c1310ba3a4968&amp;amp;map_id=4fe5fff0d41e3958&amp;amp;map_id=87551272fa78f2bd&amp;amp;map_id=2cc91e9b1b4bd533&#34;&gt;this link&lt;/a&gt; will take you to the collection of maps of Wilson&amp;rsquo;s Prom shown above.&lt;/p&gt;
&lt;h2 id=&#34;the-background&#34;&gt;The background&lt;/h2&gt;
&lt;p&gt;If you missed the start of this journey back in November last year, you might be wondering what the georeferenced maps are and where they come from. During my SLV LAB residency, I found a way of &lt;a href=&#34;https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html&#34;&gt;hooking the SLV&amp;rsquo;s digitised maps up to a tool called Allmaps&lt;/a&gt; that helps you identify points that connect historic maps to our modern coordinate system. When enough points have been identified, the historic maps can be positioned on a modern basemap. This is known as georeferencing, georectifying, or &amp;lsquo;map warping&amp;rsquo;,  as the results can often appear skewed or warped.&lt;/p&gt;
&lt;p&gt;Once I had connected things up, I invited the world (or at least the tiny part of it that follows me on social media) to help turn the SLV&amp;rsquo;s maps into data. And they did! As of today, &lt;strong&gt;1,447&lt;/strong&gt; of the SLV&amp;rsquo;s digitised maps have been georeferenced. &lt;a href=&#34;https://wragge.github.io/slv-allmaps/dashboard.html&#34;&gt;This dashboard displays current georeferencing progress&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/visualization-3.png&#34; width=&#34;600&#34; height=&#34;211&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;The total number of SLV maps georeferenced over time. It&#39;s still going up!&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;There&amp;rsquo;s still plenty more to do. If you&amp;rsquo;d like to help, &lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;the full instructions are available here&lt;/a&gt;. Georeferencing is pretty fun, so why not have a go?&lt;/p&gt;
&lt;p&gt;You can explore the current collection of georeferenced maps in a few different ways. There&amp;rsquo;s a dataset you can &lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.csv&#34;&gt;download&lt;/a&gt; or &lt;a href=&#34;https://glam-workbench.net/datasette-lite/?csv=https%3A%2F%2Fgithub.com%2Fwragge%2Fslv-allmaps%2Fblob%2Fmain%2Fgeoreferenced_maps_datasette.csv&amp;amp;install=datasette-homepage-table&amp;amp;install=datasette-json-html&amp;amp;fts=manifest_title%2Cmap_title&#34;&gt;search&lt;/a&gt; that gets updated every two hours. This data is loaded into &lt;a href=&#34;https://slv-places-481615284700.australia-southeast1.run.app/&#34;&gt;a spatial database&lt;/a&gt; that&amp;rsquo;s used by the &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Georeferenced Maps Explorer&lt;/a&gt;. As part of my recent improvements, I&amp;rsquo;ve automated this process as well, so the database should be updated with the latest additions every 24 hours.&lt;/p&gt;
&lt;p&gt;You can also search for georeferenced maps using &lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;the my place app&lt;/a&gt;. You just enter an address and my place pulls together data from a variety of sources – mixing the georeferenced maps up with parish maps, newspapers, photos, and entries from the Sands &amp;amp; MacDougall&amp;rsquo;s directories.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-geo.png&#34; width=&#34;600&#34; height=&#34;296&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Georeferenced maps in &lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;my place&lt;/a&gt; results&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&#34;the-interface&#34;&gt;The interface&lt;/h2&gt;
&lt;p&gt;The &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Georeferenced Maps Explorer&lt;/a&gt; uses &lt;a href=&#34;https://maplibre.org&#34;&gt;MapLibre&lt;/a&gt; and the &lt;a href=&#34;https://github.com/allmaps/allmaps/tree/main/packages/maplibre&#34;&gt;Allmaps MapLibre plugin&lt;/a&gt; to display the georeferenced maps. You might notice that it looks pretty similar to the &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;Newspapers Explorer&lt;/a&gt; and the &lt;a href=&#34;https://slv.wraggelabs.com/cua/&#34;&gt;CUA Browser&lt;/a&gt;, both of which use MapLibre, as well as &lt;a href=&#34;https://bulma.io&#34;&gt;Bulma&lt;/a&gt; for CSS. I&amp;rsquo;ve been trying to settle on a fairly standard set of tools that I can use to create and maintain these sorts of interfaces without too much fuss. Basically I just cut and paste a lot of stuff, then modify as needed.&lt;/p&gt;
&lt;p&gt;When you click on the basemap in the Explorer, the coordinates are sent off to the spatial database to retrieve details of georeferenced maps within 50km. The spatial database runs in &lt;a href=&#34;https://datasette.io&#34;&gt;Datasette&lt;/a&gt;, which has a built-in JSON API that I use with a set of predefined &amp;lsquo;canned&amp;rsquo; queries to pull back the data I need. The results are displayed in the right-hand column, along with square thumbnails generated by the SLV&amp;rsquo;s &lt;a href=&#34;https://iiif.io/&#34;&gt;IIIF&lt;/a&gt; service.&lt;/p&gt;
&lt;p&gt;The metadata includes distance and area measures. These are used to find and sort the results. There are two distance measures, one from your selected point to the closest boundary of a map, and the other to the centre of a map. If the point is contained within a map&amp;rsquo;s boundaries, then the &amp;lsquo;bounds&amp;rsquo; distance is zero. The search query finds maps whose closest boundaries are within 50km. Originally I sorted the results by this distance and the area of the maps. But this meant that large scale maps that included the selected point (such as maps of the whole of Victoria) appeared above nearby local maps. To make it easier to find maps within an area, I added the &amp;lsquo;centre&amp;rsquo; distance and now sort the results using that. This allows nearby maps that don&amp;rsquo;t include the current point to bubble up towards the top of the search results, above many of the large scale maps. It&amp;rsquo;s far from perfect, but I think it strikes an ok balance.&lt;/p&gt;
&lt;p&gt;The data also includes the boundaries of each map as GeoJSON. I use this to generate a MapLibre layer that contains all the boundaries as polygons. The boundaries are hidden until you hover over the corresponding search result, then the opacity of the boundary is flipped to &lt;code&gt;1&lt;/code&gt; and it magically appears.&lt;/p&gt;
&lt;p&gt;When you click on a search result, a request is fired off to &lt;a href=&#34;https://allmaps.org&#34;&gt;Allmaps&lt;/a&gt; for the full georeferencing data. The Allmaps plugin uses this to retrieve the map image from the SLV&amp;rsquo;s IIIF service and display the warped map in MapLibre.&lt;/p&gt;
&lt;p&gt;I looked around for quite a while to find a good way of changing the opacity and order of the warped maps in MapLibre. I eventually found the &lt;a href=&#34;https://github.com/wragge/maplibre-gl-layer-manager&#34;&gt;Map Libre GL Layer Manager&lt;/a&gt; which did a lot of what I wanted. I &lt;a href=&#34;https://github.com/wragge/maplibre-gl-layer-manager&#34;&gt;forked the repository&lt;/a&gt; and modified the code to get the opacity slider to work with warped map layers. Warped map layers already have a &lt;code&gt;setOpacity&lt;/code&gt; method, it was just a matter of checking for &amp;lsquo;custom&amp;rsquo; layers, then finding where the warped map was in the layer object.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4&#34;&gt;&lt;code class=&#34;language-javascript&#34; data-lang=&#34;javascript&#34;&gt;&lt;span style=&#34;color:#66d9ef&#34;&gt;if&lt;/span&gt; (&lt;span style=&#34;color:#a6e22e&#34;&gt;type&lt;/span&gt; &lt;span style=&#34;color:#f92672&#34;&gt;==&lt;/span&gt; &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;custom&amp;#34;&lt;/span&gt;) {
        &lt;span style=&#34;color:#a6e22e&#34;&gt;layer&lt;/span&gt;.&lt;span style=&#34;color:#a6e22e&#34;&gt;implementation&lt;/span&gt;.&lt;span style=&#34;color:#a6e22e&#34;&gt;setOpacity&lt;/span&gt;(&lt;span style=&#34;color:#a6e22e&#34;&gt;opacity&lt;/span&gt;);
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I also made a few cosmetic changes, such as renaming the tooltips on the reorder buttons from &amp;lsquo;move up&amp;rsquo; and &amp;lsquo;move down&amp;rsquo; to &amp;lsquo;send back&amp;rsquo; and &amp;lsquo;bring forward&amp;rsquo; – up and down just confused me.&lt;/p&gt;
&lt;p&gt;I tried for a long time to find some way of adding tooltips or popups to the warped maps that would show their details when you moved the mouse over them. I found that if you were displaying multiple maps that looked similar, such as the photomaps above, it was difficult to know which map was which. After a chat with the &lt;a href=&#34;https://allmaps.org&#34;&gt;Allmaps&lt;/a&gt; developers in their IIIF Slack channel, I realised that this approach wouldn&amp;rsquo;t work as the warped map layers don&amp;rsquo;t currently listen to mouse events. Instead I decided to add hover events to the list of results, rather than the maps, and use them to display the map boundaries as described above. This way I get the connection between the map and metadata that I wanted, as well as a useful way of previewing results.&lt;/p&gt;
&lt;p&gt;I think I&amp;rsquo;ve probably stopped fiddling with the interface for now. I hope you find it useful!&lt;/p&gt;
&lt;h2 id=&#34;the-future&#34;&gt;The future?&lt;/h2&gt;
&lt;p&gt;There&amp;rsquo;s more that I&amp;rsquo;d like to do with the georeferenced maps. In particular, I&amp;rsquo;ve been thinking about an interface with a slider that showed the changing patchwork of maps over time&amp;hellip;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related resources:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;the code for the Georeferenced Newspapers Explorer and all the other apps and sites I created during my residency is &lt;a href=&#34;https://github.com/wragge/slv-demo-apps&#34;&gt;in this GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;the code to harvest the georeferenced data from Allmaps and build the dashboard is in &lt;a href=&#34;https://github.com/wragge/slv-allmaps&#34;&gt;this GitHub repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;there&amp;rsquo;s also the &lt;a href=&#34;https://slv.wraggelabs.com&#34;&gt;full list of all the apps, code, posts, and talks&lt;/a&gt; created during my residency&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;a href=&#34;https://doi.org/10.17613/m8c1d-50882&#34;&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/doi-10.17613-m8c1d-d9b01c.svg&#34; style=&#34;border:none;width:200px&#34;&gt;&lt;/a&gt;&lt;/p&gt;
</description>
      <source:markdown>I&#39;m in the process of tying up all the documentation relating to my time as [Creative Technologist-in-Residence at the State Library of Victoria LAB](https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt/people-place-library-data-tim-sherratt). But as I was looking through [the list of outputs](https://slv.wraggelabs.com), I realised I&#39;d never written anything about the interface I created to explore georeferenced maps from the SLV collection.

I also remembered that there were a few improvements I wanted to make to the interface. So instead of spending a few hours writing up a blog post, I&#39;ve spent several days completely overhauling the [Georeferenced Maps Explorer](https://slv.wraggelabs.com/geomaps/). I&#39;m pretty happy with how it&#39;s working now. **[Have a play!](https://slv.wraggelabs.com/geomaps/)**

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/prom-maps.png&#34; width=&#34;600&#34; height=&#34;354&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Wilson&#39;s Prom made up of a patchwork of georeferenced maps and aerial photographs using the Georefrenced Maps Explorer. &lt;a href=&#34;https://slv.wraggelabs.com/geomaps/&#34;&gt;Try it now!&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;

To get started, just click on the basemap. Details of all georeferenced maps within 50km of your selected point will be displayed in the right-hand column. As you move your mouse over the list of results, the boundaries of the georeferenced maps will be displayed on the basemap. This gives you a preview of their location and size. Click on one of the results to display the georeferenced map as a layer on top of the modern basemap.

&lt;figure&gt;&lt;video src=&#34;https://cdn.uploads.micro.mov/8371/2026/video-2026-02-12-13-39-23/playlist.m3u8&#34; poster=&#34;https://cdn.uploads.micro.blog/8371/2026/frames/1679959-0-ad6b9b.jpg&#34; width=&#34;1920&#34; height=&#34;1080&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34;&gt;&lt;/video&gt;&lt;figcaption&gt;Hover over a result to see the map boundaries&lt;/figcaption&gt;&lt;/figure&gt;

You can add as many maps as you like. If your selected maps overlap, you can change the order in which they&#39;re shown. Click on the layers icon in the top left of the basemap. You&#39;ll see a list of the maps that are currently displayed. Use the arrow buttons to move a map backwards or forwards. You can also use the sliders to adjust the opacity of each map. This can make it easier to examine the relationship between maps. For example, you might want to compare the features of a historic map with those of the underlying basemap.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/photomaps.png&#34; width=&#34;600&#34; height=&#34;363&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Stitch together multiple maps like this series of seven photomaps, and change the opacity to see the features underneath&lt;/figcaption&gt;&lt;/figure&gt;

The Explorer&#39;s url updates with every selection you make, so you can bookmark or share a url to return to the same position and collection of maps. For example, [this link](https://slv.wraggelabs.com/geomaps/?lat=-38.96774552450964&amp;lon=146.39004985426214&amp;map_id=215c1310ba3a4968&amp;map_id=4fe5fff0d41e3958&amp;map_id=87551272fa78f2bd&amp;map_id=2cc91e9b1b4bd533) will take you to the collection of maps of Wilson&#39;s Prom shown above.

## The background

If you missed the start of this journey back in November last year, you might be wondering what the georeferenced maps are and where they come from. During my SLV LAB residency, I found a way of [hooking the SLV&#39;s digitised maps up to a tool called Allmaps](https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html) that helps you identify points that connect historic maps to our modern coordinate system. When enough points have been identified, the historic maps can be positioned on a modern basemap. This is known as georeferencing, georectifying, or &#39;map warping&#39;,  as the results can often appear skewed or warped.

Once I had connected things up, I invited the world (or at least the tiny part of it that follows me on social media) to help turn the SLV&#39;s maps into data. And they did! As of today, **1,447** of the SLV&#39;s digitised maps have been georeferenced. [This dashboard displays current georeferencing progress](https://wragge.github.io/slv-allmaps/dashboard.html).

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/visualization-3.png&#34; width=&#34;600&#34; height=&#34;211&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;The total number of SLV maps georeferenced over time. It&#39;s still going up!&lt;/figcaption&gt;&lt;/figure&gt;

There&#39;s still plenty more to do. If you&#39;d like to help, [the full instructions are available here](https://wragge.github.io/slv-allmaps/). Georeferencing is pretty fun, so why not have a go?

You can explore the current collection of georeferenced maps in a few different ways. There&#39;s a dataset you can [download](https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.csv) or [search](https://glam-workbench.net/datasette-lite/?csv=https%3A%2F%2Fgithub.com%2Fwragge%2Fslv-allmaps%2Fblob%2Fmain%2Fgeoreferenced_maps_datasette.csv&amp;install=datasette-homepage-table&amp;install=datasette-json-html&amp;fts=manifest_title%2Cmap_title) that gets updated every two hours. This data is loaded into [a spatial database](https://slv-places-481615284700.australia-southeast1.run.app/) that&#39;s used by the [Georeferenced Maps Explorer](https://slv.wraggelabs.com/geomaps/). As part of my recent improvements, I&#39;ve automated this process as well, so the database should be updated with the latest additions every 24 hours.

You can also search for georeferenced maps using [the my place app](https://slv.wraggelabs.com/myplace/). You just enter an address and my place pulls together data from a variety of sources – mixing the georeferenced maps up with parish maps, newspapers, photos, and entries from the Sands &amp; MacDougall&#39;s directories.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-geo.png&#34; width=&#34;600&#34; height=&#34;296&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Georeferenced maps in &lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;my place&lt;/a&gt; results&lt;/figcaption&gt;&lt;/figure&gt;

## The interface

The [Georeferenced Maps Explorer](https://slv.wraggelabs.com/geomaps/) uses [MapLibre](https://maplibre.org) and the [Allmaps MapLibre plugin](https://github.com/allmaps/allmaps/tree/main/packages/maplibre) to display the georeferenced maps. You might notice that it looks pretty similar to the [Newspapers Explorer](https://slv.wraggelabs.com/newspapers/) and the [CUA Browser](https://slv.wraggelabs.com/cua/), both of which use MapLibre, as well as [Bulma](https://bulma.io) for CSS. I&#39;ve been trying to settle on a fairly standard set of tools that I can use to create and maintain these sorts of interfaces without too much fuss. Basically I just cut and paste a lot of stuff, then modify as needed.

When you click on the basemap in the Explorer, the coordinates are sent off to the spatial database to retrieve details of georeferenced maps within 50km. The spatial database runs in [Datasette](https://datasette.io), which has a built-in JSON API that I use with a set of predefined &#39;canned&#39; queries to pull back the data I need. The results are displayed in the right-hand column, along with square thumbnails generated by the SLV&#39;s [IIIF](https://iiif.io/) service.

The metadata includes distance and area measures. These are used to find and sort the results. There are two distance measures, one from your selected point to the closest boundary of a map, and the other to the centre of a map. If the point is contained within a map&#39;s boundaries, then the &#39;bounds&#39; distance is zero. The search query finds maps whose closest boundaries are within 50km. Originally I sorted the results by this distance and the area of the maps. But this meant that large scale maps that included the selected point (such as maps of the whole of Victoria) appeared above nearby local maps. To make it easier to find maps within an area, I added the &#39;centre&#39; distance and now sort the results using that. This allows nearby maps that don&#39;t include the current point to bubble up towards the top of the search results, above many of the large scale maps. It&#39;s far from perfect, but I think it strikes an ok balance.

The data also includes the boundaries of each map as GeoJSON. I use this to generate a MapLibre layer that contains all the boundaries as polygons. The boundaries are hidden until you hover over the corresponding search result, then the opacity of the boundary is flipped to `1` and it magically appears.

When you click on a search result, a request is fired off to [Allmaps](https://allmaps.org) for the full georeferencing data. The Allmaps plugin uses this to retrieve the map image from the SLV&#39;s IIIF service and display the warped map in MapLibre.

I looked around for quite a while to find a good way of changing the opacity and order of the warped maps in MapLibre. I eventually found the [Map Libre GL Layer Manager](https://github.com/wragge/maplibre-gl-layer-manager) which did a lot of what I wanted. I [forked the repository](https://github.com/wragge/maplibre-gl-layer-manager) and modified the code to get the opacity slider to work with warped map layers. Warped map layers already have a `setOpacity` method, it was just a matter of checking for &#39;custom&#39; layers, then finding where the warped map was in the layer object.

```javascript
if (type == &#34;custom&#34;) {
        layer.implementation.setOpacity(opacity);
```
I also made a few cosmetic changes, such as renaming the tooltips on the reorder buttons from &#39;move up&#39; and &#39;move down&#39; to &#39;send back&#39; and &#39;bring forward&#39; – up and down just confused me.

I tried for a long time to find some way of adding tooltips or popups to the warped maps that would show their details when you moved the mouse over them. I found that if you were displaying multiple maps that looked similar, such as the photomaps above, it was difficult to know which map was which. After a chat with the [Allmaps](https://allmaps.org) developers in their IIIF Slack channel, I realised that this approach wouldn&#39;t work as the warped map layers don&#39;t currently listen to mouse events. Instead I decided to add hover events to the list of results, rather than the maps, and use them to display the map boundaries as described above. This way I get the connection between the map and metadata that I wanted, as well as a useful way of previewing results.

I think I&#39;ve probably stopped fiddling with the interface for now. I hope you find it useful!

## The future?

There&#39;s more that I&#39;d like to do with the georeferenced maps. In particular, I&#39;ve been thinking about an interface with a slider that showed the changing patchwork of maps over time...

**Related resources:**

- the code for the Georeferenced Newspapers Explorer and all the other apps and sites I created during my residency is [in this GitHub repository](https://github.com/wragge/slv-demo-apps)
- the code to harvest the georeferenced data from Allmaps and build the dashboard is in [this GitHub repository](https://github.com/wragge/slv-allmaps)
- there&#39;s also the [full list of all the apps, code, posts, and talks](https://slv.wraggelabs.com) created during my residency


&lt;a href=&#34;https://doi.org/10.17613/m8c1d-50882&#34;&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/doi-10.17613-m8c1d-d9b01c.svg&#34; style=&#34;border:none;width:200px&#34;&gt;&lt;/a&gt;








</source:markdown>
    </item>
    
    <item>
      <title>my place – exploring SLV collections through a street address</title>
      <link>https://updates.timsherratt.org/2026/02/02/my-place-exploring-slv-collections.html</link>
      <pubDate>Mon, 02 Feb 2026 21:43:20 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/02/02/my-place-exploring-slv-collections.html</guid>
      <description>&lt;p&gt;&lt;em&gt;&amp;lsquo;What can I find out about my house?&#39;&lt;/em&gt; My work as &lt;a href=&#34;https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt&#34;&gt;Creative Technologist-in-Residence at the SLV LAB&lt;/a&gt; was inspired by questions like this that librarians at the SLV hear every day. I wanted to explore how the Library&amp;rsquo;s place-based collections could be used to provide new entry points for discovery and navigation – entry points based not on words, but locations.&lt;/p&gt;
&lt;p&gt;At the end of my residency, I pulled all the different collections I&amp;rsquo;d been working with into a single interface – &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;. It&amp;rsquo;s not polished or complete, but I think it&amp;rsquo;s a useful starting point to think about the possibilities. You just type in an address, street name, or place name and my place shows you maps, photos, newspapers, and even extracts from the Sands &amp;amp; MacDougall directories. &lt;strong&gt;&lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;Try it now!&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-02-02-17-44-07.png&#34; width=&#34;600&#34; height=&#34;433&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;Try &lt;b&gt;&lt;i&gt;my place!&lt;/i&gt;&lt;/b&gt;&lt;/a&gt; Just enter an address in the search box.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Search results in my place are bookmarkable. So save and share your discoveries!&lt;/p&gt;
&lt;h2 id=&#34;the-collections&#34;&gt;The collections&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; draws its data from a number of different place-based collections that I&amp;rsquo;ve been working on during my residency.&lt;/p&gt;
&lt;h3 id=&#34;openstreetmap&#34;&gt;OpenStreetMap&lt;/h3&gt;
&lt;p&gt;When you enter an address in the search box, &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; looks it up in &lt;a href=&#34;https://www.openstreetmap.org&#34;&gt;OpenStreetMap&lt;/a&gt; to get its geospatial coordinates. It then places a marker and re-centres the map at the top of the app.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp.png&#34; width=&#34;600&#34; height=&#34;219&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Map centred on 149 Brunswick Street, Fitzroy&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;OpenStreetMap is also used to retrieve additional information about the suburb, including its boundaries.&lt;/p&gt;
&lt;h3 id=&#34;sands--macdougalls-directories&#34;&gt;Sands &amp;amp; MacDougall&amp;rsquo;s directories&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; queries the &lt;a href=&#34;https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html&#34;&gt;full-text searchable version of Sands &amp;amp; Mac&lt;/a&gt; for addresses. Results will vary based on the OCR quality and the nature of query, but it can give you a potted history of who has lived in your house. The search results are displayed in chronological order, and include an &lt;a href=&#34;https://updates.timsherratt.org/2025/11/16/some-sands-mac-tweaks-thanks.html&#34;&gt;image snippet&lt;/a&gt; showing the actual printed entry as well as the text content and metadata.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-sandm.png&#34; width=&#34;600&#34; height=&#34;349&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Occupants of 149 Brunswick Street, Fitzroy from 1875 to 1925&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h3 id=&#34;committee-for-urban-action-photographs&#34;&gt;Committee for Urban Action photographs&lt;/h3&gt;
&lt;p&gt;If you enter a full street address, &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; will search &lt;a href=&#34;https://updates.timsherratt.org/2026/01/29/geolocating-photos-from-the-slvs.html&#34;&gt;the CUA collection&lt;/a&gt; for photos associated with the segment of road that includes the current address. It then displays the individual images from any matching photosets.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-cua.png&#34; width=&#34;600&#34; height=&#34;432&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Photographs from CUA of the currently selected road&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Otherwise &lt;em&gt;&lt;strong&gt;my space&lt;/strong&gt;&lt;/em&gt; will look for CUA photos that are near the current location, and display a randomly-selected image from each photoset.&lt;/p&gt;
&lt;h3 id=&#34;georeferenced-maps&#34;&gt;Georeferenced maps&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; searches through &lt;a href=&#34;https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html&#34;&gt;digitised maps from the SLV collection that have been georeferenced by the public&lt;/a&gt;. It finds maps that either intersect with the currently selected location, or are nearby.&lt;/p&gt;
&lt;p&gt;If you enter a full street address, the first 6 georeferenced maps will be positioned on a modern basemap with a marker indicating the currently selected point. This means you can see your address on a historical map. The number of georeferenced maps that can be displayed in this way is determined by the browser – so I&amp;rsquo;ve limited it to 6 to be safe.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-geo.png&#34; width=&#34;600&#34; height=&#34;296&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Georeferenced maps positioned on a modern basemap, showing the location of the currently selected address&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h3 id=&#34;parish-maps&#34;&gt;Parish maps&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; searches through &lt;a href=&#34;https://updates.timsherratt.org/2025/10/06/creating-bounding-boxes-for-parish.html&#34;&gt;parish maps in the SLV collection that have geospatial coordinates or approximate bounding boxes&lt;/a&gt;. It finds maps that either intersect with the currently selected location, or are nearby.&lt;/p&gt;
&lt;h3 id=&#34;newspapers&#34;&gt;Newspapers&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; searches through &lt;a href=&#34;https://updates.timsherratt.org/2025/12/16/exploring-victorian-newspapers.html&#34;&gt;my dataset of newspapers in the SLV collection&lt;/a&gt; that have a place of publication documented in the &amp;lsquo;Place newspaper published&amp;rsquo; metadata field. It finds newspapers that are either associated with the current suburb/town, or a nearby suburb/town. This includes digitised and non-digitised titles. Digitised titles include a link to Trove.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-newspapers.png&#34; width=&#34;600&#34; height=&#34;293&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Newspapers from the SLV collection published in Fitzroy&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h3 id=&#34;photographs&#34;&gt;Photographs&lt;/h3&gt;
&lt;p&gt;I thought it would be cool to include a few photographs of the current suburb or town. To do this, I downloaded a list of place names from VicNames, then used the place names to &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_place_images.ipynb&#34;&gt;search the SLV catalogue for photographs with relevant subject headings&lt;/a&gt;. A random selection of the harvested images is displayed in &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-images.png&#34; width=&#34;600&#34; height=&#34;234&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;A few images of Fitzroy, displayed alongside a map of Fitzroy&#39;s current boundaries using data from OpenStreetMap&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&#34;the-interface&#34;&gt;The interface&lt;/h2&gt;
&lt;p&gt;The interface is pretty simple. You type an address in the box and hit enter. If the geocoding process finds multiple matches, it&amp;rsquo;ll give you a list to choose from. Once the location is found, a marker is added and the main map re-centres. Then related resources are displayed below the map.&lt;/p&gt;
&lt;p&gt;As you scroll down through the results you gradually zoom out from your initial starting point. This is reflected in the four bands or layers used to group resources: &amp;lsquo;my house&amp;rsquo;, &amp;lsquo;my street&amp;rsquo;, &amp;lsquo;my suburb&amp;rsquo;, and &amp;lsquo;nearby&amp;rsquo;. Each band contains a mix of resources from different collections.&lt;/p&gt;
&lt;p&gt;When I started working on &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;, I was thinking about a project from around 2010 called &lt;a href=&#34;https://wraggelabs.com/info/history-wall/&#34;&gt;The History Wall&lt;/a&gt;. Like &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;, The History Wall pulled many different types of resources together into a rich exploratory interface. As you scrolled through The History Wall you moved through time, with randomly selected items appearing from a range of sources including Trove newspapers, the ADB, and museum collections.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/history-wall.jpg&#34; width=&#34;600&#34; height=&#34;505&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;A version of The History Wall created for the National Museum of Australia&#39;s &#39;Irish in Australia&#39; exhibition&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;I originally thought I&amp;rsquo;d inject some of the same randomness into &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;, but I was worried it might just get too confusing. I thought it was important to keep the relationship between the starting point and the resources in focus even as you zoomed out. So my visual metaphor shifted to something more like a blast radius map, or a stratigraphic diagram, that displayed distinct groups and layers as you moved beyond the baseline. My limited CSS skills couldn&amp;rsquo;t make the vision in my head a reality, but there are lots of headings and colours instead to highlight the transitions!&lt;/p&gt;
&lt;p&gt;The actual mix of groups and layers displayed depends on the nature of your query. If you&amp;rsquo;ve entered a complete street address, and there are results for that address in Sands &amp;amp; Mac, then you&amp;rsquo;ll see &amp;lsquo;my house&amp;rsquo;, &amp;lsquo;my suburb&amp;rsquo;, and &amp;lsquo;nearby&amp;rsquo;. If you&amp;rsquo;ve only entered a suburb or town, or your street address can&amp;rsquo;t be found, you&amp;rsquo;ll see two layers starting with &amp;lsquo;my suburb&amp;rsquo;.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s an overview of what you might expect to see.&lt;/p&gt;
&lt;h3 id=&#34;my-house&#34;&gt;my house&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;Sands &amp;amp; MacDougall extracts (text search on full address)&lt;/li&gt;
&lt;li&gt;georeferenced maps (search for maps that contain the base point)&lt;/li&gt;
&lt;li&gt;parish maps (search for maps that contain the base point)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;my-street&#34;&gt;my street&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;CUA photos (search for matching street identifiers)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;if there&amp;rsquo;s no &amp;lsquo;my house&amp;rsquo; layer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Sands &amp;amp; MacDougall extracts (text search on street name and suburb)&lt;/li&gt;
&lt;li&gt;georeferenced maps (search for intersections between maps and street)&lt;/li&gt;
&lt;li&gt;parish maps (search for intersections between maps and street)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;my-suburbtown&#34;&gt;my suburb/town&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;suburb boundaries from OSM&lt;/li&gt;
&lt;li&gt;images (search for suburb name in metadata)&lt;/li&gt;
&lt;li&gt;newspapers (search for suburb name in metadata)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;if there&amp;rsquo;s no &amp;lsquo;my house&amp;rsquo; or &amp;lsquo;my street&amp;rsquo; layer:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;georeferenced maps (search for intersections between maps and suburb boundaries)&lt;/li&gt;
&lt;li&gt;parish maps (search for intersections between maps and suburb boundaries)&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id=&#34;nearby&#34;&gt;nearby&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;CUA photos (search for photosets within 5km of the base point, filtered to remove current street)&lt;/li&gt;
&lt;li&gt;georeferenced maps (search for maps within 10km of base point, ordered by distance, max of 24 displayed)&lt;/li&gt;
&lt;li&gt;parish maps (search for maps within 10km of base point, ordered by distance, max of 24 displayed)&lt;/li&gt;
&lt;li&gt;newspapers (search for newspapers within 100km of base point, ordered by distance, max of 24 displayed)&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id=&#34;the-data&#34;&gt;The data&lt;/h2&gt;
&lt;p&gt;Most of the data used in &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; is stored in two SQLite databases – &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;one for Sands &amp;amp; Mac&lt;/a&gt;, and &lt;a href=&#34;https://slv-places-481615284700.australia-southeast1.run.app/&#34;&gt;the other for CUA, georeferenced maps, parish maps, and newspapers&lt;/a&gt;. The metadata for the collection images is stored in &lt;a href=&#34;https://raw.githubusercontent.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/refs/heads/main/place_images.json&#34;&gt;a JSON file&lt;/a&gt; that is directly loaded by the interface.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve published the SQLite databases online using &lt;a href=&#34;https://datasette.io&#34;&gt;Datasette&lt;/a&gt; and &lt;a href=&#34;https://www.gaia-gis.it/fossil/libspatialite/index&#34;&gt;Spatialite&lt;/a&gt;. Spatialite makes it possible to find geospatial features that intersect, or are near, a given point. For example, you could find maps that include a specific set of coordinates.&lt;/p&gt;
&lt;p&gt;Datasette has the ability to create &lt;a href=&#34;https://docs.datasette.io/en/stable/sql_queries.html#canned-queries&#34;&gt;&amp;lsquo;canned queries&amp;rsquo;&lt;/a&gt; that feed url parameters into pre-defined SQL queries. This coupled with Datasette&amp;rsquo;s &lt;a href=&#34;https://docs.datasette.io/en/stable/json_api.html&#34;&gt;built-in JSON API&lt;/a&gt; makes it possible to construct query urls in &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; and use them to retrieve JSON results data from my databases.&lt;/p&gt;
&lt;p&gt;When you enter an address in &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;, multiple queries are fired off to find intersecting or nearby resources. For example, this url finds georeferenced maps within 10km of a point at the centre of Fitzroy: &lt;a href=&#34;https://slv-places-481615284700.australia-southeast1.run.app/georeferenced_maps/maps_from_wkt.json?wkt=POINT(144.977468%20-37.803143)&#34;&gt;slv-places-481615284700.australia-southeast1.run.app/georefere&amp;hellip;&lt;/a&gt;&amp;amp;distance=10000&amp;amp;_shape=array.&lt;/p&gt;
&lt;p&gt;In the case of Sands &amp;amp; Mac, &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; uses a canned query that runs a full-text search across the OCRd content of a volume. Suburb names are often abbreviated in Sands &amp;amp; Mac, so the app first runs a query to find possible abbreviations, then adds them into the main query to inject a bit of fuzziness. This is repeated for all 24 digitised volumes.&lt;/p&gt;
&lt;p&gt;Once the metadata is retrieved from the databases, images are loaded from the SLV&amp;rsquo;s IIIF service.&lt;/p&gt;
&lt;h2 id=&#34;next-steps&#34;&gt;Next steps?&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;m not sure how much more work I&amp;rsquo;ll do on &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt;, but there are a few things I&amp;rsquo;d like to try. In particular, I&amp;rsquo;d like to help the user understand more about what data is being presented, or not presented, and why. Not all digitised maps have been georeferenced, not all parish maps have coordinates, street numbers have changed, and the OCR in Sands &amp;amp; Mac varies in quality. &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; can only present a sample – a gesture towards the wealth of material available from the SLV. I feel that message needs to be made more explicit. Though I&amp;rsquo;m not sure how without overloading the interface.&lt;/p&gt;
&lt;p&gt;There are additional data sources I&amp;rsquo;d like to play around with. &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; already includes some code to query &lt;a href=&#34;https://www.wikidata.org/&#34;&gt;Wikidata&lt;/a&gt; for more information about a suburb. But I haven&amp;rsquo;t had a chance to do anything with it. I&amp;rsquo;d like to be able to provide additional contextual information from outside the SLV, such as electoral boundaries, populations, even election results. It would also be fun to display sightings of plants and animals from the &lt;a href=&#34;https://www.ala.org.au&#34;&gt;Atlas of Living Australia&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;What can I find out about my house? It would be great if &lt;em&gt;&lt;strong&gt;my place&lt;/strong&gt;&lt;/em&gt; could answer that question by taking the user on an open ended journey through our cultural, historical, and environmental landscape.&lt;/p&gt;
</description>
      <source:markdown>*&#39;What can I find out about my house?&#39;* My work as [Creative Technologist-in-Residence at the SLV LAB](https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt) was inspired by questions like this that librarians at the SLV hear every day. I wanted to explore how the Library&#39;s place-based collections could be used to provide new entry points for discovery and navigation – entry points based not on words, but locations.

At the end of my residency, I pulled all the different collections I&#39;d been working with into a single interface – ***my place***. It&#39;s not polished or complete, but I think it&#39;s a useful starting point to think about the possibilities. You just type in an address, street name, or place name and my place shows you maps, photos, newspapers, and even extracts from the Sands &amp; MacDougall directories. **[Try it now!](https://slv.wraggelabs.com/myplace/)**

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2026-02-02-17-44-07.png&#34; width=&#34;600&#34; height=&#34;433&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;Try &lt;b&gt;&lt;i&gt;my place!&lt;/i&gt;&lt;/b&gt;&lt;/a&gt; Just enter an address in the search box.&lt;/figcaption&gt;&lt;/figure&gt;

Search results in my place are bookmarkable. So save and share your discoveries!

## The collections

***my place*** draws its data from a number of different place-based collections that I&#39;ve been working on during my residency.

### OpenStreetMap

When you enter an address in the search box, ***my place*** looks it up in [OpenStreetMap](https://www.openstreetmap.org) to get its geospatial coordinates. It then places a marker and re-centres the map at the top of the app.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp.png&#34; width=&#34;600&#34; height=&#34;219&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Map centred on 149 Brunswick Street, Fitzroy&lt;/figcaption&gt;&lt;/figure&gt;

OpenStreetMap is also used to retrieve additional information about the suburb, including its boundaries.

### Sands &amp; MacDougall&#39;s directories

***my place*** queries the [full-text searchable version of Sands &amp; Mac](https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html) for addresses. Results will vary based on the OCR quality and the nature of query, but it can give you a potted history of who has lived in your house. The search results are displayed in chronological order, and include an [image snippet](https://updates.timsherratt.org/2025/11/16/some-sands-mac-tweaks-thanks.html) showing the actual printed entry as well as the text content and metadata. 

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-sandm.png&#34; width=&#34;600&#34; height=&#34;349&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Occupants of 149 Brunswick Street, Fitzroy from 1875 to 1925&lt;/figcaption&gt;&lt;/figure&gt;

### Committee for Urban Action photographs

If you enter a full street address, ***my place*** will search [the CUA collection](https://updates.timsherratt.org/2026/01/29/geolocating-photos-from-the-slvs.html) for photos associated with the segment of road that includes the current address. It then displays the individual images from any matching photosets.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-cua.png&#34; width=&#34;600&#34; height=&#34;432&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Photographs from CUA of the currently selected road&lt;/figcaption&gt;&lt;/figure&gt;

Otherwise ***my space*** will look for CUA photos that are near the current location, and display a randomly-selected image from each photoset.

### Georeferenced maps

***my place*** searches through [digitised maps from the SLV collection that have been georeferenced by the public](https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html). It finds maps that either intersect with the currently selected location, or are nearby.

If you enter a full street address, the first 6 georeferenced maps will be positioned on a modern basemap with a marker indicating the currently selected point. This means you can see your address on a historical map. The number of georeferenced maps that can be displayed in this way is determined by the browser – so I&#39;ve limited it to 6 to be safe.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-geo.png&#34; width=&#34;600&#34; height=&#34;296&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Georeferenced maps positioned on a modern basemap, showing the location of the currently selected address&lt;/figcaption&gt;&lt;/figure&gt;

### Parish maps

***my place*** searches through [parish maps in the SLV collection that have geospatial coordinates or approximate bounding boxes](https://updates.timsherratt.org/2025/10/06/creating-bounding-boxes-for-parish.html). It finds maps that either intersect with the currently selected location, or are nearby.

### Newspapers

***my place*** searches through [my dataset of newspapers in the SLV collection](https://updates.timsherratt.org/2025/12/16/exploring-victorian-newspapers.html) that have a place of publication documented in the &#39;Place newspaper published&#39; metadata field. It finds newspapers that are either associated with the current suburb/town, or a nearby suburb/town. This includes digitised and non-digitised titles. Digitised titles include a link to Trove.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-newspapers.png&#34; width=&#34;600&#34; height=&#34;293&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Newspapers from the SLV collection published in Fitzroy&lt;/figcaption&gt;&lt;/figure&gt;

### Photographs

I thought it would be cool to include a few photographs of the current suburb or town. To do this, I downloaded a list of place names from VicNames, then used the place names to [search the SLV catalogue for photographs with relevant subject headings](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_place_images.ipynb). A random selection of the harvested images is displayed in ***my place***.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/mp-images.png&#34; width=&#34;600&#34; height=&#34;234&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;A few images of Fitzroy, displayed alongside a map of Fitzroy&#39;s current boundaries using data from OpenStreetMap&lt;/figcaption&gt;&lt;/figure&gt;

## The interface

The interface is pretty simple. You type an address in the box and hit enter. If the geocoding process finds multiple matches, it&#39;ll give you a list to choose from. Once the location is found, a marker is added and the main map re-centres. Then related resources are displayed below the map.

As you scroll down through the results you gradually zoom out from your initial starting point. This is reflected in the four bands or layers used to group resources: &#39;my house&#39;, &#39;my street&#39;, &#39;my suburb&#39;, and &#39;nearby&#39;. Each band contains a mix of resources from different collections.

When I started working on ***my place***, I was thinking about a project from around 2010 called [The History Wall](https://wraggelabs.com/info/history-wall/). Like ***my place***, The History Wall pulled many different types of resources together into a rich exploratory interface. As you scrolled through The History Wall you moved through time, with randomly selected items appearing from a range of sources including Trove newspapers, the ADB, and museum collections.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/history-wall.jpg&#34; width=&#34;600&#34; height=&#34;505&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;A version of The History Wall created for the National Museum of Australia&#39;s &#39;Irish in Australia&#39; exhibition&lt;/figcaption&gt;&lt;/figure&gt;

I originally thought I&#39;d inject some of the same randomness into ***my place***, but I was worried it might just get too confusing. I thought it was important to keep the relationship between the starting point and the resources in focus even as you zoomed out. So my visual metaphor shifted to something more like a blast radius map, or a stratigraphic diagram, that displayed distinct groups and layers as you moved beyond the baseline. My limited CSS skills couldn&#39;t make the vision in my head a reality, but there are lots of headings and colours instead to highlight the transitions!

The actual mix of groups and layers displayed depends on the nature of your query. If you&#39;ve entered a complete street address, and there are results for that address in Sands &amp; Mac, then you&#39;ll see &#39;my house&#39;, &#39;my suburb&#39;, and &#39;nearby&#39;. If you&#39;ve only entered a suburb or town, or your street address can&#39;t be found, you&#39;ll see two layers starting with &#39;my suburb&#39;.

Here&#39;s an overview of what you might expect to see.

### my house

- Sands &amp; MacDougall extracts (text search on full address)
- georeferenced maps (search for maps that contain the base point)
- parish maps (search for maps that contain the base point)

### my street 

- CUA photos (search for matching street identifiers)

if there&#39;s no &#39;my house&#39; layer:

- Sands &amp; MacDougall extracts (text search on street name and suburb)
- georeferenced maps (search for intersections between maps and street)
- parish maps (search for intersections between maps and street)

### my suburb/town

- suburb boundaries from OSM
- images (search for suburb name in metadata)
- newspapers (search for suburb name in metadata)

if there&#39;s no &#39;my house&#39; or &#39;my street&#39; layer:

- georeferenced maps (search for intersections between maps and suburb boundaries)
- parish maps (search for intersections between maps and suburb boundaries)

### nearby

- CUA photos (search for photosets within 5km of the base point, filtered to remove current street)
- georeferenced maps (search for maps within 10km of base point, ordered by distance, max of 24 displayed)
- parish maps (search for maps within 10km of base point, ordered by distance, max of 24 displayed)
- newspapers (search for newspapers within 100km of base point, ordered by distance, max of 24 displayed)

## The data

Most of the data used in ***my place*** is stored in two SQLite databases – [one for Sands &amp; Mac](https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/), and [the other for CUA, georeferenced maps, parish maps, and newspapers](https://slv-places-481615284700.australia-southeast1.run.app/). The metadata for the collection images is stored in [a JSON file](https://raw.githubusercontent.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/refs/heads/main/place_images.json) that is directly loaded by the interface.

I&#39;ve published the SQLite databases online using [Datasette](https://datasette.io) and [Spatialite](https://www.gaia-gis.it/fossil/libspatialite/index). Spatialite makes it possible to find geospatial features that intersect, or are near, a given point. For example, you could find maps that include a specific set of coordinates.

Datasette has the ability to create [&#39;canned queries&#39;](https://docs.datasette.io/en/stable/sql_queries.html#canned-queries) that feed url parameters into pre-defined SQL queries. This coupled with Datasette&#39;s [built-in JSON API](https://docs.datasette.io/en/stable/json_api.html) makes it possible to construct query urls in ***my place*** and use them to retrieve JSON results data from my databases.

When you enter an address in ***my place***, multiple queries are fired off to find intersecting or nearby resources. For example, this url finds georeferenced maps within 10km of a point at the centre of Fitzroy: [slv-places-481615284700.australia-southeast1.run.app/georefere...](https://slv-places-481615284700.australia-southeast1.run.app/georeferenced_maps/maps_from_wkt.json?wkt=POINT(144.977468%20-37.803143))&amp;distance=10000&amp;_shape=array.

In the case of Sands &amp; Mac, ***my place*** uses a canned query that runs a full-text search across the OCRd content of a volume. Suburb names are often abbreviated in Sands &amp; Mac, so the app first runs a query to find possible abbreviations, then adds them into the main query to inject a bit of fuzziness. This is repeated for all 24 digitised volumes.

Once the metadata is retrieved from the databases, images are loaded from the SLV&#39;s IIIF service.

## Next steps?

I&#39;m not sure how much more work I&#39;ll do on ***my place***, but there are a few things I&#39;d like to try. In particular, I&#39;d like to help the user understand more about what data is being presented, or not presented, and why. Not all digitised maps have been georeferenced, not all parish maps have coordinates, street numbers have changed, and the OCR in Sands &amp; Mac varies in quality. ***my place*** can only present a sample – a gesture towards the wealth of material available from the SLV. I feel that message needs to be made more explicit. Though I&#39;m not sure how without overloading the interface.

There are additional data sources I&#39;d like to play around with. ***my place*** already includes some code to query [Wikidata](https://www.wikidata.org/) for more information about a suburb. But I haven&#39;t had a chance to do anything with it. I&#39;d like to be able to provide additional contextual information from outside the SLV, such as electoral boundaries, populations, even election results. It would also be fun to display sightings of plants and animals from the [Atlas of Living Australia](https://www.ala.org.au).

What can I find out about my house? It would be great if ***my place*** could answer that question by taking the user on an open ended journey through our cultural, historical, and environmental landscape.





</source:markdown>
    </item>
    
    <item>
      <title>Geolocating photos from the SLV&#39;s Committee for Urban Action collection</title>
      <link>https://updates.timsherratt.org/2026/01/29/geolocating-photos-from-the-slvs.html</link>
      <pubDate>Thu, 29 Jan 2026 17:08:06 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2026/01/29/geolocating-photos-from-the-slvs.html</guid>
      <description>&lt;p&gt;Concerned about the loss of built heritage in the 1970s, the Committee for Urban Action photographed streetscapes across urban and regional Victoria. They compiled a remarkable collection of photographs that is &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;amp;collectionId=81271917420007636&#34;&gt;now being digitised by the State Library of Victoria&lt;/a&gt;. More than 20,000 images are already available online!&lt;/p&gt;
&lt;p&gt;The CUA worked systematically, capturing photos street by street, and recording the locations of each set of photographs. This information is used to prepare the title attached to each photo as it&amp;rsquo;s uploaded to the SLV catalogue. In general, titles include the name of the road where the photo was taken, the name of the suburb or town, and the names of two intersecting roads that define the boundaries of the current road segment. They can also tell you which side of the road the photo was taken on. For example, the title &lt;code&gt;Gore Street, Fitzroy, from Gertrude Street to Webb Street - east side&lt;/code&gt; tells us the photo was taken on the east side of Gore Street, Fitzroy between the intersections with Gertrude Street and Webb Street.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-slv-viewer.png&#34; width=&#34;600&#34; height=&#34;547&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Photos from Gore Street, Fitzroy &lt;a href=&#34;https://viewer.slv.vic.gov.au/?entity=IE7489506&amp;mode=browse&#34;&gt;displayed in the SLV image viewer&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;It&amp;rsquo;s great to have this sort of structured information linking photos to specific locations, but to navigate through the collection &lt;em&gt;in space&lt;/em&gt; we need more. We need to link each photo to a set of geospatial coordinates by mapping each road segment. That was the challenge I took on as part of &lt;a href=&#34;https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt&#34;&gt;my residency in the SLV LAB&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;When I started working on the collection I wasn&amp;rsquo;t really sure what was possible. I had to learn a lot, and ended up revising my processes multiple times as I got deeper into the data. But my aim was always to create some sort of map-based interface, that would allow users to click on a street and see any associated CUA photos. It&amp;rsquo;s still a bit buggy and incomplete, but here it is – &lt;a href=&#34;https://slv.wraggelabs.com/cua/&#34;&gt;&lt;strong&gt;explore the CUA collection street by street&lt;/strong&gt;&lt;/a&gt;!&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-browser.png&#34; width=&#34;600&#34; height=&#34;513&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Gore Street, Fitzroy &lt;a href=&#34;https://slv.wraggelabs.com/cua/?photoset=gore-street-fitzroy-gertrude-street-webb-street&#34;&gt;in the new CUA Browser&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&#34;the-process&#34;&gt;The process&lt;/h2&gt;
&lt;p&gt;My basic plan was to find the intersections using &lt;a href=&#34;https://www.openstreetmap.org/&#34;&gt;OpenStreetMap&lt;/a&gt;, then extract geospatial information about the segment of road between the two intersections. This involved much trial and error, but eventually I ended up with a process that:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;parsed each item title to try and extract the names of the main road, the suburb, and the two intersecting roads&lt;/li&gt;
&lt;li&gt;queried &lt;a href=&#34;https://nominatim.org&#34;&gt;Nominatim&lt;/a&gt; for the suburb bounding box&lt;/li&gt;
&lt;li&gt;for each intersecting road, queried OSM to find a node at, or around, its intersection with the main road, within the suburb bounding box&lt;/li&gt;
&lt;li&gt;created a new bounding box from the coordinates of the two intersections&lt;/li&gt;
&lt;li&gt;queried OSM for the main road within this bounding box&lt;/li&gt;
&lt;li&gt;extracted the coordinates of the main road segment, removing any points outside of the bounding box&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;There&amp;rsquo;s more details below and in these notebooks: &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua_finding_intersections.ipynb&#34;&gt;cua_finding_intersections.ipynb&lt;/a&gt; and &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua_data_processing.ipynb&#34;&gt;cua_data_processing.ipynb&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;finding-intersections&#34;&gt;Finding intersections&lt;/h2&gt;
&lt;p&gt;As described, the title of each photograph generally includes 4 pieces of information: the road, suburb, intersecting roads, and side. My plan was to find the intersections first to get the limits of the road segment. This is possible thanks to the awesome &lt;a href=&#34;https://www.openstreetmap.org/&#34;&gt;OpenStreetMap&lt;/a&gt; and its &lt;a href=&#34;https://wiki.openstreetmap.org/wiki/Overpass_API&#34;&gt;Overpass API&lt;/a&gt;. It took me a while to get my head around the Overpass query language, but there are lots of &lt;a href=&#34;https://wiki.openstreetmap.org/wiki/Overpass_API/Overpass_API_by_Example&#34;&gt;useful examples online&lt;/a&gt;. The query to find the intersection between Gore Street and Gertrude Street in Fitzroy looks like this:&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;[bbox:-37.8089071,144.9732006,-37.7929130,144.9851430];
way[&#39;highway&#39;][name=&amp;quot;Gore Street&amp;quot;];
node(w)-&amp;gt;.n1;
way[&#39;highway&#39;][name=&amp;quot;Gertrude Street&amp;quot;];
node(w)-&amp;gt;.n2;
node.n1.n2; 
out body;
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;You can &lt;a href=&#34;https://overpass-turbo.eu/s/2jq8&#34;&gt;try it out&lt;/a&gt; using Overpass Turbo&amp;rsquo;s web interface.&lt;/p&gt;
&lt;p&gt;In OpenStreetMap, linear features, such as roads or rivers, are represented as &lt;a href=&#34;https://wiki.openstreetmap.org/wiki/Way&#34;&gt;&lt;code&gt;ways&lt;/code&gt;&lt;/a&gt;. Each way is made up of a series of &lt;code&gt;nodes&lt;/code&gt; or points with geospatial coordinates. Every way and node has its own unique identifier. Tags can be added to features to describe what type of things they are.&lt;/p&gt;
&lt;p&gt;The query above looks for &lt;code&gt;ways&lt;/code&gt; named &amp;lsquo;Gore Street&amp;rsquo; and &amp;lsquo;Gertrude Street&amp;rsquo; that are tagged as &lt;code&gt;highway&lt;/code&gt; (a &lt;code&gt;highway&lt;/code&gt; in OpenStreetMap is any road-like feature including things bike paths and foot trails).&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code class=&#34;language-way[&#39;highway&#39;][name=&#34;Gore&#34; data-lang=&#34;way[&#39;highway&#39;][name=&#34;Gore&#34;&gt;way[&#39;highway&#39;][name=&amp;quot;Gore Street&amp;quot;];
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;It then extracts the nodes that make up each way and looks to see if there are any nodes in common between the two ways.  A node shared between two ways indicates an intersection.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;node(w)-&amp;gt;.n2;
node.n1.n2; 
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The query is limited using a bounding box that encloses the suburb of Fitzroy. This avoids false positives and keeps down the query load.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;[bbox:-37.8089071,144.9732006,-37.7929130,144.9851430];
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The JSON result of this query gives as the latitude and longitude of the node at the intersection of the two roads.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;{
  &amp;quot;version&amp;quot;: 0.6,
  &amp;quot;generator&amp;quot;: &amp;quot;Overpass API 0.7.62.10 2d4cfc48&amp;quot;,
  &amp;quot;osm3s&amp;quot;: {
    &amp;quot;timestamp_osm_base&amp;quot;: &amp;quot;2026-01-27T03:11:45Z&amp;quot;,
    &amp;quot;copyright&amp;quot;: &amp;quot;The data included in this document is from www.openstreetmap.org. The data is made available under ODbL.&amp;quot;
  },
  &amp;quot;elements&amp;quot;: [

{
  &amp;quot;type&amp;quot;: &amp;quot;node&amp;quot;,
  &amp;quot;id&amp;quot;: 224750459,
  &amp;quot;lat&amp;quot;: -37.8062302,
  &amp;quot;lon&amp;quot;: 144.9817848
}

  ]
}
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;After a bit of testing, I found this worked pretty well, except for roundabouts&amp;hellip; In OpenStreetMap, roads don&amp;rsquo;t actually cross roundabouts – they end on one side, then begin anew on the other side. In cases like this, looking for shared nodes doesn&amp;rsquo;t work. Instead you have to look to see if the two roads have nodes that are less than a given distance apart. The query is similar to the one above, but uses &lt;code&gt;around&lt;/code&gt; when comparing the nodes. In this case I&amp;rsquo;m looking for nodes that are within 20 metres of each other.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;node(w.w2)(around.w1:20);
&lt;/code&gt;&lt;/pre&gt;&lt;h2 id=&#34;finding-road-segments&#34;&gt;Finding road segments&lt;/h2&gt;
&lt;p&gt;Once I had the coordinates of the two intersections, I could look for the segment of road between between them. To do this I created a bounding box using the coordinates of the intersections, and then searched for ways by name within that defined area.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s important to note that there&amp;rsquo;s no one-to-one correspondence between roads and OSM ways. A single road might be represented in OSM as a series of separate, but connected, ways. For example, at a roundabout, or where a road divides, new ways might have been created to document the change. This means that when we query OSM for details of a road we often get back information about multiple ways. Some of these might be things like bike paths which we can filter using tags, but often they&amp;rsquo;ll be sections of the road that we want. For example, this query for Gore Street, within the bounds of its intersections with Gertrude Street and Webb Street, returns details of two ways.&lt;/p&gt;
&lt;pre tabindex=&#34;0&#34;&gt;&lt;code&gt;way[&amp;quot;highway&amp;quot;~&amp;quot;^(trunk|primary|secondary|tertiary|unclassified|residential|service|track|pedestrian|living_street)$&amp;quot;][name=&amp;quot;Gore Street&amp;quot;](-37.8062302,144.98128480000003,-37.8040076,144.9826827);
out body;
&amp;gt;;
out body;
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;You can &lt;a href=&#34;https://overpass-turbo.eu/s/2jqi&#34;&gt;view the result&lt;/a&gt; in Overpass Turbo.&lt;/p&gt;
&lt;p&gt;However, that doesn&amp;rsquo;t mean that the full extent of both ways is contained within the bounding box, just that some of the nodes of both ways are inside. Because of this, I filtered the results from all the ways and only kept nodes whose coordinates were within the desired region.&lt;/p&gt;
&lt;h2 id=&#34;problems-finding-intersections&#34;&gt;Problems finding intersections&lt;/h2&gt;
&lt;p&gt;The method described above works pretty well, and once I understood enough about the Overpass API to get out actual paths that I could display on a map, I fed all of the CUA photos through a script and got useful data for more than 80% of them.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2025-09-27-17-34-36.png&#34; width=&#34;600&#34; height=&#34;652&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;One of my early tests.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Then I spent a &lt;em&gt;lot of&lt;/em&gt; time trying to understand where the remainder were failing.&lt;/p&gt;
&lt;p&gt;Some of them failed because the titles were missing information, or were formatted in a way I didn&amp;rsquo;t expect. For example, instead of a second intersecting road, some titles just said &amp;lsquo;to end&amp;rsquo;. This makes perfect sense to a human looking at a map, but it&amp;rsquo;s difficult to handle programmatically.&lt;/p&gt;
&lt;p&gt;Some photos either recorded the wrong suburb, or the boundaries of the suburb had moved since the photos were taken. For example, many of the photos described as being from Eaglehawk are now in California Gully.&lt;/p&gt;
&lt;p&gt;Similarly, some road names were wrong either because of documentation errors, or because the names have changed over time. There are also some variations in the way OSM records road names – in particular, I found that roads with hyphenated names sometimes had spaces around the hyphen and sometimes didn&amp;rsquo;t. There were also a couple of cases where names weren&amp;rsquo;t attached to the corresponding road segment in OSM, but I was able to edit these in OSM directly.&lt;/p&gt;
&lt;p&gt;Other roads had multiple names, or change names along their path. I mean, what&amp;rsquo;s going on with Brunswick Street and St Georges Road in Fitzroy? Country towns seemed most prone to this – a highway might become &amp;lsquo;Main Road&amp;rsquo; within the town boundaries, or the order of hyphenated places in road names might change. I found one road in Clunes that had four different names within the space of a few hundred metres.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/clunes.png&#34; width=&#34;582&#34; height=&#34;1002&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;One road, four names!&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Finally, the routes of some roads had changed – intersections no longer intersected, roads were closed, or new parks had popped up to split a road in two.&lt;/p&gt;
&lt;p&gt;My processing script logged the titles I couldn&amp;rsquo;t locate and I worked through the list manually, trying to identify what each problem was. I suppose there&amp;rsquo;s two ways I could&amp;rsquo;ve handled these problems – building more fuzziness into the process to check for things like alternative names, or by compiling a list of &amp;lsquo;corrected&amp;rsquo; titles. I started off using the first approach, but as I worked through more and more anomalies, the checking logic became very complicated and inefficient. Just think about the knots you can tie yourself in trying to handle a title where the suburb is wrong and the main road changes names in between intersections.&lt;/p&gt;
&lt;p&gt;I refactored the code multiple times, but it&amp;rsquo;s still pretty messy. In the end I created a list of &amp;lsquo;corrected&amp;rsquo; titles as well, so it was a bit of a hybrid approach. I suspect I could have saved myself a lot of pain if I&amp;rsquo;d reversed the process – compiling &amp;lsquo;corrected&amp;rsquo; titles first, then adapting the logic as patterns emerged.&lt;/p&gt;
&lt;p&gt;There are still some photos I haven&amp;rsquo;t located. In some cases I just don&amp;rsquo;t have enough information. In others I need to manually record coordinates or way ids to feed into the process, and I haven&amp;rsquo;t worked out the best way to do this yet. You can see the titles that I&amp;rsquo;ve haven&amp;rsquo;t geolocated yet in the files: &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-not-found.txt&#34;&gt;&lt;code&gt;cua-not-found.txt&lt;/code&gt;&lt;/a&gt; and &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-not-parsed.txt&#34;&gt;&lt;code&gt;cua-not-parsed.txt&lt;/code&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;In total, 18,603 out of 20,644 photos have been geolocated. That&amp;rsquo;s over 90%!&lt;/p&gt;
&lt;h2 id=&#34;assembling-the-data&#34;&gt;Assembling the data&lt;/h2&gt;
&lt;p&gt;I processed the data in a couple of phases to get it in the shape I wanted.&lt;/p&gt;
&lt;p&gt;The first step was to group all the photos by title, so I could link each group to its location. But remember that titles often record which &lt;em&gt;side&lt;/em&gt; of the road a photo was taken on. To bring all sides of a road segment together into a single group, I created a key from a normalised/slugified version of the title with the side value removed. I used this key to save information about each side within the same group.&lt;/p&gt;
&lt;p&gt;I ended up with a dataset with this sort of structure (a truncated example):&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4&#34;&gt;&lt;code class=&#34;language-json&#34; data-lang=&#34;json&#34;&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;iffla-street-south-melbourne-coventry-street-normanby-street&amp;#34;&lt;/span&gt;&lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;:&lt;/span&gt;
    {
        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;title&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;Iffla Street, South Melbourne, from Coventry Street to Normanby Street&amp;#34;&lt;/span&gt;,
        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;sides&amp;#34;&lt;/span&gt;:
        {
            &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;east side&amp;#34;&lt;/span&gt;:
            {
                &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;title&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;Iffla Street, South Melbourne, from Coventry Street to Normanby Street - east side.&amp;#34;&lt;/span&gt;,
                &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;images&amp;#34;&lt;/span&gt;:
                [
                    {
                        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;ie_id&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;IE20321667&amp;#34;&lt;/span&gt;,
                        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;alma_id&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#ae81ff&#34;&gt;9939649155207636&lt;/span&gt;
                    }
                    &lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;more&lt;/span&gt; &lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;photos...&lt;/span&gt;
                ]
            },
            &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;west side&amp;#34;&lt;/span&gt;:
            {
                &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;title&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;Iffla Street, South Melbourne, from Normanby Street to Coventry Street - west side.&amp;#34;&lt;/span&gt;,
                &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;images&amp;#34;&lt;/span&gt;:
                [
                    {
                        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;ie_id&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;IE20320072&amp;#34;&lt;/span&gt;,
                        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;alma_id&amp;#34;&lt;/span&gt;: &lt;span style=&#34;color:#ae81ff&#34;&gt;9939655629407636&lt;/span&gt;
                    },
					&lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;more&lt;/span&gt; &lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;photos...&lt;/span&gt;
                ]
            }
        },
        &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;ways&amp;#34;&lt;/span&gt;:
        {
            &lt;span style=&#34;color:#f92672&#34;&gt;&amp;#34;27631235&amp;#34;&lt;/span&gt;:
            [
                [
                    &lt;span style=&#34;color:#ae81ff&#34;&gt;144.9503379&lt;/span&gt;,
                    &lt;span style=&#34;color:#ae81ff&#34;&gt;-37.835322&lt;/span&gt;
                ],
                &lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;more&lt;/span&gt; &lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;points...&lt;/span&gt;
            ]
        }
    }&lt;span style=&#34;color:#960050;background-color:#1e0010&#34;&gt;,&lt;/span&gt;
		
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;You can see how the sides and matching ways have been brought together under the key value.&lt;/p&gt;
&lt;p&gt;This structure was useful for grouping and processing the data, but to create a map interface I needed to bring the geospatial information to the surface. The first version of the interface used one big GeoJSON file in which the features were &lt;a href=&#34;https://geocrystal.github.io/geojson/GeoJSON/MultiLineString.html&#34;&gt;MultiLineStrings&lt;/a&gt; created from the paths of each road segment. The photo data was saved in the properties of each GeoJSON feature.&lt;/p&gt;
&lt;p&gt;It sort of worked. The roads with photos were highlighted, and clicking on the roads displayed the photos. It was only when I changed the opacity of the lines that I realised that, in many cases, different road segments were being piled on top of each other. When the lines were opaque these piles were invisible, but add a bit of transparency and you could see that some lines were darker than others. Clicking on the lines only displayed the top layer, so some groups of photos were effectively invisible.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2025-12-01-14-09-47.png&#34; width=&#34;503&#34; height=&#34;339&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Version one of the interface showing how the colour of the highlighted roads varied once I decreased the opacity.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Why did this happen? I&amp;rsquo;d wrongly assumed that each segment of road would only have one group of photos associated with it. But it&amp;rsquo;s not hard to find cases where this is not true. Consider Moor Street, Fitzroy, between Nicholson Street and Brunswick Street. On the north side, there is a single group of photos that document the buildings between Nicholson Street and Brunswick Street. However, on the south side there&amp;rsquo;s two groups of photos. One covers the section between Nicholson Street and Fitzroy Street, the other covers Fitzroy Street to Brunswick Street. One section of road, three groups of photos&amp;hellip;&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-multiple.png&#34; width=&#34;600&#34; height=&#34;478&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Moor Street, Fitzroy, between Nicholson Street and Brunswick Street, in the new CUA Browser, showing the three photosets associated with the one section of road.&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;To make these layered groups more easily accessible through the interface I had to change the way the data was organised – separating the GeoJSON from the photosets so that multiple photosets could be associated with a single geospatial feature. I decided to create a GeoJSON feature for every OSM way in the dataset. However, I needed to prune the way&amp;rsquo;s coordinates to only include those that were part of the CUA road segments. To do this, I saved all the way data when I found the road segments. Then in the second processing phase, I grouped the way coordinates associated with the road segments and compared this list to the full way path. Any coordinate in the way path that wasn&amp;rsquo;t in the road segments was removed. It seems unnecessarily complex, but I wanted to make sure that only the parts of roads associated with photos were highlighted in the interface.&lt;/p&gt;
&lt;p&gt;The result was two data files. The first, &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-ways.geojson&#34;&gt;&lt;code&gt;cua-ways.geojson&lt;/code&gt;&lt;/a&gt;, contains the pruned way paths and their OSM identifiers. The second, &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-photos.json&#34;&gt;&lt;code&gt;cua-photos.json&lt;/code&gt;&lt;/a&gt;, contains information about each photo set, including the sides, photos, paths, and associated way identifiers. The datasets are linked by the way identifiers.&lt;/p&gt;
&lt;h2 id=&#34;constructing-the-interface&#34;&gt;Constructing the interface&lt;/h2&gt;
&lt;p&gt;My plan for the interface was pretty simple. There&amp;rsquo;d be a map on which all the road segments associated with CUA photos were highlighted. Clicking on a highlighted section would show the photos. I wanted to display the photos as if you were scanning the streetscape, so I decided to put them all side-by-side in a gallery that scrolled horizontally.&lt;/p&gt;
&lt;p&gt;The first version used Leaflet to display the maps and, as noted above, had some problems where there were multiple photosets associated with a segment of road.&lt;/p&gt;
&lt;p&gt;For the &lt;a href=&#34;https://slv.wraggelabs.com/cua/&#34;&gt;second version&lt;/a&gt; I decided to switch to &lt;a href=&#34;https://maplibre.org&#34;&gt;MapLibre&lt;/a&gt; because it seems a bit more active and up-to-date. I&amp;rsquo;d already used MapLibre in the &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;SLV Newspapers Explorer&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The interface first loads the  &lt;code&gt;cua-ways.geojson&lt;/code&gt; file to highlight the relevant roads. When you click on one of the roads, the way id is passed to a function that looks for associated photo sets in the &lt;code&gt;cua-photos.json&lt;/code&gt; data. If there&amp;rsquo;s only one linked photoset, then the photos are displayed. However, if there&amp;rsquo;s more than one linked photoset, they&amp;rsquo;re displayed as a list. The user then selects from the list to display the related photos.&lt;/p&gt;
&lt;p&gt;A couple of other things happen when you click on a way or select a photoset:  the colour of the selected road segment changes, and the browser url is updated with the way or photoset identifier. You can bookmark or share these urls to go directly to a specific road or photoset. There&amp;rsquo;s also a button to reverse the order of the images – they scroll left to right, but sometimes they seem to have been photographed right to left.&lt;/p&gt;
&lt;h2 id=&#34;more-information-and-links&#34;&gt;More information and links&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;a href=&#34;https://slv.wraggelabs.com/cua/&#34;&gt;CUA Browser&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;CUA data is also used in the &lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;my place&lt;/a&gt; app&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;CUA code and data is in the &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency&#34;&gt;geo-maps-residency&lt;/a&gt; repository&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Code for the interface is in the &lt;a href=&#34;https://github.com/wragge/slv-demo-apps&#34;&gt;slv-demo-apps&lt;/a&gt; repository&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;all the outcomes of my SLV residency are listed on the &lt;a href=&#34;https://slv.wraggelabs.com&#34;&gt;Experiments with the State Library of Victoria’s collections&lt;/a&gt; page&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
</description>
      <source:markdown>Concerned about the loss of built heritage in the 1970s, the Committee for Urban Action photographed streetscapes across urban and regional Victoria. They compiled a remarkable collection of photographs that is [now being digitised by the State Library of Victoria](https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;collectionId=81271917420007636). More than 20,000 images are already available online!

The CUA worked systematically, capturing photos street by street, and recording the locations of each set of photographs. This information is used to prepare the title attached to each photo as it&#39;s uploaded to the SLV catalogue. In general, titles include the name of the road where the photo was taken, the name of the suburb or town, and the names of two intersecting roads that define the boundaries of the current road segment. They can also tell you which side of the road the photo was taken on. For example, the title `Gore Street, Fitzroy, from Gertrude Street to Webb Street - east side` tells us the photo was taken on the east side of Gore Street, Fitzroy between the intersections with Gertrude Street and Webb Street.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-slv-viewer.png&#34; width=&#34;600&#34; height=&#34;547&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Photos from Gore Street, Fitzroy &lt;a href=&#34;https://viewer.slv.vic.gov.au/?entity=IE7489506&amp;mode=browse&#34;&gt;displayed in the SLV image viewer&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;

It&#39;s great to have this sort of structured information linking photos to specific locations, but to navigate through the collection *in space* we need more. We need to link each photo to a set of geospatial coordinates by mapping each road segment. That was the challenge I took on as part of [my residency in the SLV LAB](https://lab.slv.vic.gov.au/experiments/my-place-tim-sherratt).

When I started working on the collection I wasn&#39;t really sure what was possible. I had to learn a lot, and ended up revising my processes multiple times as I got deeper into the data. But my aim was always to create some sort of map-based interface, that would allow users to click on a street and see any associated CUA photos. It&#39;s still a bit buggy and incomplete, but here it is – [**explore the CUA collection street by street**](https://slv.wraggelabs.com/cua/)!

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-browser.png&#34; width=&#34;600&#34; height=&#34;513&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Gore Street, Fitzroy &lt;a href=&#34;https://slv.wraggelabs.com/cua/?photoset=gore-street-fitzroy-gertrude-street-webb-street&#34;&gt;in the new CUA Browser&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;

## The process

My basic plan was to find the intersections using [OpenStreetMap](https://www.openstreetmap.org/), then extract geospatial information about the segment of road between the two intersections. This involved much trial and error, but eventually I ended up with a process that:

- parsed each item title to try and extract the names of the main road, the suburb, and the two intersecting roads
- queried [Nominatim](https://nominatim.org) for the suburb bounding box
- for each intersecting road, queried OSM to find a node at, or around, its intersection with the main road, within the suburb bounding box
- created a new bounding box from the coordinates of the two intersections
- queried OSM for the main road within this bounding box
- extracted the coordinates of the main road segment, removing any points outside of the bounding box

There&#39;s more details below and in these notebooks: [cua_finding_intersections.ipynb](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua_finding_intersections.ipynb) and [cua_data_processing.ipynb](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua_data_processing.ipynb).

## Finding intersections

As described, the title of each photograph generally includes 4 pieces of information: the road, suburb, intersecting roads, and side. My plan was to find the intersections first to get the limits of the road segment. This is possible thanks to the awesome [OpenStreetMap](https://www.openstreetmap.org/) and its [Overpass API](https://wiki.openstreetmap.org/wiki/Overpass_API). It took me a while to get my head around the Overpass query language, but there are lots of [useful examples online](https://wiki.openstreetmap.org/wiki/Overpass_API/Overpass_API_by_Example). The query to find the intersection between Gore Street and Gertrude Street in Fitzroy looks like this:

```
[bbox:-37.8089071,144.9732006,-37.7929130,144.9851430];
way[&#39;highway&#39;][name=&#34;Gore Street&#34;];
node(w)-&gt;.n1;
way[&#39;highway&#39;][name=&#34;Gertrude Street&#34;];
node(w)-&gt;.n2;
node.n1.n2; 
out body;
```
You can [try it out](https://overpass-turbo.eu/s/2jq8) using Overpass Turbo&#39;s web interface.

In OpenStreetMap, linear features, such as roads or rivers, are represented as [`ways`](https://wiki.openstreetmap.org/wiki/Way). Each way is made up of a series of `nodes` or points with geospatial coordinates. Every way and node has its own unique identifier. Tags can be added to features to describe what type of things they are. 

The query above looks for `ways` named &#39;Gore Street&#39; and &#39;Gertrude Street&#39; that are tagged as `highway` (a `highway` in OpenStreetMap is any road-like feature including things bike paths and foot trails).

```way[&#39;highway&#39;][name=&#34;Gore Street&#34;];
way[&#39;highway&#39;][name=&#34;Gore Street&#34;];
```
It then extracts the nodes that make up each way and looks to see if there are any nodes in common between the two ways.  A node shared between two ways indicates an intersection.

```
node(w)-&gt;.n2;
node.n1.n2; 
```
The query is limited using a bounding box that encloses the suburb of Fitzroy. This avoids false positives and keeps down the query load.

```
[bbox:-37.8089071,144.9732006,-37.7929130,144.9851430];
```
The JSON result of this query gives as the latitude and longitude of the node at the intersection of the two roads.

```
{
  &#34;version&#34;: 0.6,
  &#34;generator&#34;: &#34;Overpass API 0.7.62.10 2d4cfc48&#34;,
  &#34;osm3s&#34;: {
    &#34;timestamp_osm_base&#34;: &#34;2026-01-27T03:11:45Z&#34;,
    &#34;copyright&#34;: &#34;The data included in this document is from www.openstreetmap.org. The data is made available under ODbL.&#34;
  },
  &#34;elements&#34;: [

{
  &#34;type&#34;: &#34;node&#34;,
  &#34;id&#34;: 224750459,
  &#34;lat&#34;: -37.8062302,
  &#34;lon&#34;: 144.9817848
}

  ]
}
```
After a bit of testing, I found this worked pretty well, except for roundabouts... In OpenStreetMap, roads don&#39;t actually cross roundabouts – they end on one side, then begin anew on the other side. In cases like this, looking for shared nodes doesn&#39;t work. Instead you have to look to see if the two roads have nodes that are less than a given distance apart. The query is similar to the one above, but uses `around` when comparing the nodes. In this case I&#39;m looking for nodes that are within 20 metres of each other.

```
node(w.w2)(around.w1:20);
```
## Finding road segments

Once I had the coordinates of the two intersections, I could look for the segment of road between between them. To do this I created a bounding box using the coordinates of the intersections, and then searched for ways by name within that defined area.

It&#39;s important to note that there&#39;s no one-to-one correspondence between roads and OSM ways. A single road might be represented in OSM as a series of separate, but connected, ways. For example, at a roundabout, or where a road divides, new ways might have been created to document the change. This means that when we query OSM for details of a road we often get back information about multiple ways. Some of these might be things like bike paths which we can filter using tags, but often they&#39;ll be sections of the road that we want. For example, this query for Gore Street, within the bounds of its intersections with Gertrude Street and Webb Street, returns details of two ways.

```
way[&#34;highway&#34;~&#34;^(trunk|primary|secondary|tertiary|unclassified|residential|service|track|pedestrian|living_street)$&#34;][name=&#34;Gore Street&#34;](-37.8062302,144.98128480000003,-37.8040076,144.9826827);
out body;
&gt;;
out body;
```
You can [view the result](https://overpass-turbo.eu/s/2jqi) in Overpass Turbo.

However, that doesn&#39;t mean that the full extent of both ways is contained within the bounding box, just that some of the nodes of both ways are inside. Because of this, I filtered the results from all the ways and only kept nodes whose coordinates were within the desired region.

## Problems finding intersections

The method described above works pretty well, and once I understood enough about the Overpass API to get out actual paths that I could display on a map, I fed all of the CUA photos through a script and got useful data for more than 80% of them. 

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2025-09-27-17-34-36.png&#34; width=&#34;600&#34; height=&#34;652&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;One of my early tests.&lt;/figcaption&gt;&lt;/figure&gt;

Then I spent a *lot of* time trying to understand where the remainder were failing.

Some of them failed because the titles were missing information, or were formatted in a way I didn&#39;t expect. For example, instead of a second intersecting road, some titles just said &#39;to end&#39;. This makes perfect sense to a human looking at a map, but it&#39;s difficult to handle programmatically.

Some photos either recorded the wrong suburb, or the boundaries of the suburb had moved since the photos were taken. For example, many of the photos described as being from Eaglehawk are now in California Gully.

Similarly, some road names were wrong either because of documentation errors, or because the names have changed over time. There are also some variations in the way OSM records road names – in particular, I found that roads with hyphenated names sometimes had spaces around the hyphen and sometimes didn&#39;t. There were also a couple of cases where names weren&#39;t attached to the corresponding road segment in OSM, but I was able to edit these in OSM directly.

Other roads had multiple names, or change names along their path. I mean, what&#39;s going on with Brunswick Street and St Georges Road in Fitzroy? Country towns seemed most prone to this – a highway might become &#39;Main Road&#39; within the town boundaries, or the order of hyphenated places in road names might change. I found one road in Clunes that had four different names within the space of a few hundred metres. 

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/clunes.png&#34; width=&#34;582&#34; height=&#34;1002&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;One road, four names!&lt;/figcaption&gt;&lt;/figure&gt;

Finally, the routes of some roads had changed – intersections no longer intersected, roads were closed, or new parks had popped up to split a road in two.

My processing script logged the titles I couldn&#39;t locate and I worked through the list manually, trying to identify what each problem was. I suppose there&#39;s two ways I could&#39;ve handled these problems – building more fuzziness into the process to check for things like alternative names, or by compiling a list of &#39;corrected&#39; titles. I started off using the first approach, but as I worked through more and more anomalies, the checking logic became very complicated and inefficient. Just think about the knots you can tie yourself in trying to handle a title where the suburb is wrong and the main road changes names in between intersections.

I refactored the code multiple times, but it&#39;s still pretty messy. In the end I created a list of &#39;corrected&#39; titles as well, so it was a bit of a hybrid approach. I suspect I could have saved myself a lot of pain if I&#39;d reversed the process – compiling &#39;corrected&#39; titles first, then adapting the logic as patterns emerged.

There are still some photos I haven&#39;t located. In some cases I just don&#39;t have enough information. In others I need to manually record coordinates or way ids to feed into the process, and I haven&#39;t worked out the best way to do this yet. You can see the titles that I&#39;ve haven&#39;t geolocated yet in the files: [`cua-not-found.txt`](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-not-found.txt) and [`cua-not-parsed.txt`](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-not-parsed.txt).

In total, 18,603 out of 20,644 photos have been geolocated. That&#39;s over 90%!

## Assembling the data

I processed the data in a couple of phases to get it in the shape I wanted.

The first step was to group all the photos by title, so I could link each group to its location. But remember that titles often record which *side* of the road a photo was taken on. To bring all sides of a road segment together into a single group, I created a key from a normalised/slugified version of the title with the side value removed. I used this key to save information about each side within the same group.

I ended up with a dataset with this sort of structure (a truncated example):

```json
&#34;iffla-street-south-melbourne-coventry-street-normanby-street&#34;:
    {
        &#34;title&#34;: &#34;Iffla Street, South Melbourne, from Coventry Street to Normanby Street&#34;,
        &#34;sides&#34;:
        {
            &#34;east side&#34;:
            {
                &#34;title&#34;: &#34;Iffla Street, South Melbourne, from Coventry Street to Normanby Street - east side.&#34;,
                &#34;images&#34;:
                [
                    {
                        &#34;ie_id&#34;: &#34;IE20321667&#34;,
                        &#34;alma_id&#34;: 9939649155207636
                    }
                    more photos...
                ]
            },
            &#34;west side&#34;:
            {
                &#34;title&#34;: &#34;Iffla Street, South Melbourne, from Normanby Street to Coventry Street - west side.&#34;,
                &#34;images&#34;:
                [
                    {
                        &#34;ie_id&#34;: &#34;IE20320072&#34;,
                        &#34;alma_id&#34;: 9939655629407636
                    },
					more photos...
                ]
            }
        },
        &#34;ways&#34;:
        {
            &#34;27631235&#34;:
            [
                [
                    144.9503379,
                    -37.835322
                ],
                more points...
            ]
        }
    },
		
```
You can see how the sides and matching ways have been brought together under the key value.

This structure was useful for grouping and processing the data, but to create a map interface I needed to bring the geospatial information to the surface. The first version of the interface used one big GeoJSON file in which the features were [MultiLineStrings](https://geocrystal.github.io/geojson/GeoJSON/MultiLineString.html) created from the paths of each road segment. The photo data was saved in the properties of each GeoJSON feature.

It sort of worked. The roads with photos were highlighted, and clicking on the roads displayed the photos. It was only when I changed the opacity of the lines that I realised that, in many cases, different road segments were being piled on top of each other. When the lines were opaque these piles were invisible, but add a bit of transparency and you could see that some lines were darker than others. Clicking on the lines only displayed the top layer, so some groups of photos were effectively invisible.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/screenshot-from-2025-12-01-14-09-47.png&#34; width=&#34;503&#34; height=&#34;339&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Version one of the interface showing how the colour of the highlighted roads varied once I decreased the opacity.&lt;/figcaption&gt;&lt;/figure&gt;

Why did this happen? I&#39;d wrongly assumed that each segment of road would only have one group of photos associated with it. But it&#39;s not hard to find cases where this is not true. Consider Moor Street, Fitzroy, between Nicholson Street and Brunswick Street. On the north side, there is a single group of photos that document the buildings between Nicholson Street and Brunswick Street. However, on the south side there&#39;s two groups of photos. One covers the section between Nicholson Street and Fitzroy Street, the other covers Fitzroy Street to Brunswick Street. One section of road, three groups of photos...

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2026/cua-multiple.png&#34; width=&#34;600&#34; height=&#34;478&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Moor Street, Fitzroy, between Nicholson Street and Brunswick Street, in the new CUA Browser, showing the three photosets associated with the one section of road.&lt;/figcaption&gt;&lt;/figure&gt;

To make these layered groups more easily accessible through the interface I had to change the way the data was organised – separating the GeoJSON from the photosets so that multiple photosets could be associated with a single geospatial feature. I decided to create a GeoJSON feature for every OSM way in the dataset. However, I needed to prune the way&#39;s coordinates to only include those that were part of the CUA road segments. To do this, I saved all the way data when I found the road segments. Then in the second processing phase, I grouped the way coordinates associated with the road segments and compared this list to the full way path. Any coordinate in the way path that wasn&#39;t in the road segments was removed. It seems unnecessarily complex, but I wanted to make sure that only the parts of roads associated with photos were highlighted in the interface.

The result was two data files. The first, [`cua-ways.geojson`](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-ways.geojson), contains the pruned way paths and their OSM identifiers. The second, [`cua-photos.json`](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/cua-photos.json), contains information about each photo set, including the sides, photos, paths, and associated way identifiers. The datasets are linked by the way identifiers.

## Constructing the interface

My plan for the interface was pretty simple. There&#39;d be a map on which all the road segments associated with CUA photos were highlighted. Clicking on a highlighted section would show the photos. I wanted to display the photos as if you were scanning the streetscape, so I decided to put them all side-by-side in a gallery that scrolled horizontally.

The first version used Leaflet to display the maps and, as noted above, had some problems where there were multiple photosets associated with a segment of road.

For the [second version](https://slv.wraggelabs.com/cua/) I decided to switch to [MapLibre](https://maplibre.org) because it seems a bit more active and up-to-date. I&#39;d already used MapLibre in the [SLV Newspapers Explorer](https://slv.wraggelabs.com/newspapers/).

The interface first loads the  `cua-ways.geojson` file to highlight the relevant roads. When you click on one of the roads, the way id is passed to a function that looks for associated photo sets in the `cua-photos.json` data. If there&#39;s only one linked photoset, then the photos are displayed. However, if there&#39;s more than one linked photoset, they&#39;re displayed as a list. The user then selects from the list to display the related photos.

A couple of other things happen when you click on a way or select a photoset:  the colour of the selected road segment changes, and the browser url is updated with the way or photoset identifier. You can bookmark or share these urls to go directly to a specific road or photoset. There&#39;s also a button to reverse the order of the images – they scroll left to right, but sometimes they seem to have been photographed right to left.

## More information and links

- [CUA Browser](https://slv.wraggelabs.com/cua/)

- CUA data is also used in the [my place](https://slv.wraggelabs.com/myplace/) app

- CUA code and data is in the [geo-maps-residency](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency) repository
- Code for the interface is in the [slv-demo-apps](https://github.com/wragge/slv-demo-apps) repository
- all the outcomes of my SLV residency are listed on the [Experiments with the State Library of Victoria’s collections](https://slv.wraggelabs.com) page
</source:markdown>
    </item>
    
    <item>
      <title>Goodbye 2025! A brief summary of the highlights and lowlights…</title>
      <link>https://updates.timsherratt.org/2025/12/31/goodbye-a-brief-summary-of.html</link>
      <pubDate>Wed, 31 Dec 2025 16:14:14 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/12/31/goodbye-a-brief-summary-of.html</guid>
      <description>&lt;p&gt;My 2025 started badly and ended well. In the first few months of the year, battles with the gatekeepers at Trove sent me spiralling into a pretty dark place. But by year’s end I was having fun, working with the wonderful people at the State Library of Victoria. In between I caught up on some overdue project maintenance. Most of this is documented in the &lt;a href=&#34;https://updates.timsherratt.org/archive/&#34;&gt;37 blog posts I wrote this year&lt;/a&gt;, but here’s a quick summary.&lt;/p&gt;&lt;h2&gt;Highlights&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;From September to December, I was Creative Technologist-in-Residence at the SLV LAB, exploring ways of opening up the Library’s place-based collections. There’s still a few things to finish off, but &lt;a href=&#34;https://slv.wraggelabs.com/&#34;&gt;here’s list of the results so far&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;As part of my SLV work, I created a &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;fully searchable version&lt;/a&gt; of the 24 volumes of Sands &amp;amp; MacDougall directories digitised by the Library. This followed the pattern I’d used for 54 volumes of the &lt;a href=&#34;https://glam-workbench.net/trove-journals/nsw-post-office-directories/&#34;&gt;NSW Post Office Directories&lt;/a&gt;, 44 volumes of the &lt;a href=&#34;https://glam-workbench.net/trove-journals/sydney-telephone-directories/&#34;&gt;Sydney Telephone Directories&lt;/a&gt;, and 54 volumes of the &lt;a href=&#34;https://glam-workbench.net/tasmanian-post-office-directories/&#34;&gt;Tasmanian Post Office Directories&lt;/a&gt;. So there’s now 176 volumes from the 1880s to the 1950s that can be easily explored for people and places— and all free to use of course.&lt;/li&gt;&lt;li&gt;In April, I added a &lt;a href=&#34;https://updates.timsherratt.org/2025/04/30/new-prov-section-added-to.html&#34;&gt;new section to the GLAM Workbench&lt;/a&gt; documenting the Public Record Office Victoria’s collection API. I also used the API to create a &lt;a href=&#34;https://updates.timsherratt.org/2025/04/10/using-the-public-record-office.html&#34;&gt;Data Dashboard that provides an overview of PROV’s collection&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;Also in April, I updated the GLAM Name Index Search &lt;a href=&#34;https://updates.timsherratt.org/2025/04/09/more-than-million-rows-of.html&#34;&gt;to include an additional 6 million records from PROV&lt;/a&gt;. In total, the GLAM Name Index now includes more that 12 million records in 293 datasets from 10 Australian GLAM organisations — another free resource for Australian researchers.&lt;/li&gt;&lt;li&gt;In July I undertook some overdue maintenance on a variety of old apps and projects. In the process, I &lt;a href=&#34;https://updates.timsherratt.org/2025/07/09/the-rebirth-of-wragge-labs.html&#34;&gt;resurrected my old Wragge Labs domain and created a showcase&lt;/a&gt; of many of the websites, apps and experiments I’ve worked on over the past 30 years.&lt;/li&gt;&lt;li&gt;I was particularly pleased &lt;a href=&#34;https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html&#34;&gt;to get &lt;em&gt;The Future of the Past&lt;/em&gt; working again&lt;/a&gt;, so once more you can create fridge magnet poetry from an odd collection of words harvested from Trove newspapers! I built FOTP back in 2012 when I was the Harold White Fellow at the NLA. Also this year I finally got around to &lt;a href=&#34;https://updates.timsherratt.org/2025/06/30/mining-for-meanings.html&#34;&gt;transcribing my Harold White Lecture&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;In June I wrote a &lt;a href=&#34;https://updates.timsherratt.org/2025/06/05/glam-workbench-preprint-for-building.html&#34;&gt;short piece on the GLAM Workbench&lt;/a&gt; for the forthcoming publication &lt;em&gt;Building User-Friendly Toolkits and Platforms for Digital Humanities&lt;/em&gt;. I think it provides a useful summary of what the GLAM Workbench is, and what I’d like it to be. I also wrote up the &lt;a href=&#34;https://updates.timsherratt.org/2025/06/19/a-brief-and-biased-history.html&#34;&gt;short but glorious history of Trove Twitter bots&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;Lowlights&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;Saying goodbye to 15 years of work on Trove&lt;/a&gt;. It still hurts. And I still miss resources such as @TroveNewsBot and the Trove API Console which ran happily for more than a decade before being killed without warning by the NLA.&lt;/li&gt;&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html&#34;&gt;Saying goodbye to 17 years of work on the National Archives of Australia’s collections&lt;/a&gt;. This will be the first New Year’s Day in a decade when I haven’t updated my &lt;a href=&#34;https://updates.timsherratt.org/2025/02/05/ten-years-of-data-the.html&#34;&gt;harvest of files with the access status of closed&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;Next year&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;In 2026, I’m looking forward to starting work on the &lt;a href=&#34;https://ardc.edu.au/project/reusable-and-accessible-public-interest-documents-rapid/&#34;&gt;RAPID project&lt;/a&gt;, building on the work I’ve done on Commonwealth Hansard over the years to create new examples and documentation.&lt;/li&gt;&lt;li&gt;I’m honoured to be giving the closing keynote at the &lt;a href=&#34;https://www.glamlabs.io/events/glam-labs-futures-26&#34;&gt;GLAM Labs Futures conference&lt;/a&gt; in Edinburgh in June — hoping we can pull together the funds to get there in person!&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;How you can help&lt;/h2&gt;&lt;p&gt;Much of my work is unfunded, and keeping resources such as the GLAM Name Index running costs real money. I’ve been very grateful for the support of my GitHub sponsors over past years. Their contributions help cover a substantial proportion of my cloud hosting costs. But bidding farewell to Twitter and Trove has had an impact on my sponsorship income. If you use or value the things I build to help researchers make use of GLAM collections, you might like to &lt;a href=&#34;https://github.com/sponsors/wragge&#34;&gt;sponsor me on GitHub&lt;/a&gt;, or &lt;a href=&#34;https://www.buymeacoffee.com/wragge&#34;&gt;Buy Me a Coffee&lt;/a&gt;. All contributions are greatly appreciated!&lt;/p&gt;&lt;p&gt;If you can’t afford a financial contribution, there are other ways you can help!&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Let me know how you’re using my stuff! A bit of positive feedback does wonders when my enthusiasm is flagging. You can find my contact details at &lt;a href=&#34;https://timsherratt.au/&#34;&gt;timsherratt.au&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;Tell others how you use my stuff! Getting information about resources out to those who might benefit is really hard, so your help would be greatly appreciated.&lt;/li&gt;&lt;li&gt;The GLAM Workbench describes a few other ways &lt;a href=&#34;https://glam-workbench.net/get-involved/&#34;&gt;you can get involved&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Goodbye 2025!&lt;/p&gt;
</description>
      <source:markdown>&lt;p&gt;My 2025 started badly and ended well. In the first few months of the year, battles with the gatekeepers at Trove sent me spiralling into a pretty dark place. But by year’s end I was having fun, working with the wonderful people at the State Library of Victoria. In between I caught up on some overdue project maintenance. Most of this is documented in the &lt;a href=&#34;https://updates.timsherratt.org/archive/&#34;&gt;37 blog posts I wrote this year&lt;/a&gt;, but here’s a quick summary.&lt;/p&gt;&lt;h2&gt;Highlights&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;From September to December, I was Creative Technologist-in-Residence at the SLV LAB, exploring ways of opening up the Library’s place-based collections. There’s still a few things to finish off, but &lt;a href=&#34;https://slv.wraggelabs.com/&#34;&gt;here’s list of the results so far&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;As part of my SLV work, I created a &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;fully searchable version&lt;/a&gt; of the 24 volumes of Sands &amp;amp; MacDougall directories digitised by the Library. This followed the pattern I’d used for 54 volumes of the &lt;a href=&#34;https://glam-workbench.net/trove-journals/nsw-post-office-directories/&#34;&gt;NSW Post Office Directories&lt;/a&gt;, 44 volumes of the &lt;a href=&#34;https://glam-workbench.net/trove-journals/sydney-telephone-directories/&#34;&gt;Sydney Telephone Directories&lt;/a&gt;, and 54 volumes of the &lt;a href=&#34;https://glam-workbench.net/tasmanian-post-office-directories/&#34;&gt;Tasmanian Post Office Directories&lt;/a&gt;. So there’s now 176 volumes from the 1880s to the 1950s that can be easily explored for people and places— and all free to use of course.&lt;/li&gt;&lt;li&gt;In April, I added a &lt;a href=&#34;https://updates.timsherratt.org/2025/04/30/new-prov-section-added-to.html&#34;&gt;new section to the GLAM Workbench&lt;/a&gt; documenting the Public Record Office Victoria’s collection API. I also used the API to create a &lt;a href=&#34;https://updates.timsherratt.org/2025/04/10/using-the-public-record-office.html&#34;&gt;Data Dashboard that provides an overview of PROV’s collection&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;Also in April, I updated the GLAM Name Index Search &lt;a href=&#34;https://updates.timsherratt.org/2025/04/09/more-than-million-rows-of.html&#34;&gt;to include an additional 6 million records from PROV&lt;/a&gt;. In total, the GLAM Name Index now includes more that 12 million records in 293 datasets from 10 Australian GLAM organisations — another free resource for Australian researchers.&lt;/li&gt;&lt;li&gt;In July I undertook some overdue maintenance on a variety of old apps and projects. In the process, I &lt;a href=&#34;https://updates.timsherratt.org/2025/07/09/the-rebirth-of-wragge-labs.html&#34;&gt;resurrected my old Wragge Labs domain and created a showcase&lt;/a&gt; of many of the websites, apps and experiments I’ve worked on over the past 30 years.&lt;/li&gt;&lt;li&gt;I was particularly pleased &lt;a href=&#34;https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html&#34;&gt;to get &lt;em&gt;The Future of the Past&lt;/em&gt; working again&lt;/a&gt;, so once more you can create fridge magnet poetry from an odd collection of words harvested from Trove newspapers! I built FOTP back in 2012 when I was the Harold White Fellow at the NLA. Also this year I finally got around to &lt;a href=&#34;https://updates.timsherratt.org/2025/06/30/mining-for-meanings.html&#34;&gt;transcribing my Harold White Lecture&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;In June I wrote a &lt;a href=&#34;https://updates.timsherratt.org/2025/06/05/glam-workbench-preprint-for-building.html&#34;&gt;short piece on the GLAM Workbench&lt;/a&gt; for the forthcoming publication &lt;em&gt;Building User-Friendly Toolkits and Platforms for Digital Humanities&lt;/em&gt;. I think it provides a useful summary of what the GLAM Workbench is, and what I’d like it to be. I also wrote up the &lt;a href=&#34;https://updates.timsherratt.org/2025/06/19/a-brief-and-biased-history.html&#34;&gt;short but glorious history of Trove Twitter bots&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;Lowlights&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;Saying goodbye to 15 years of work on Trove&lt;/a&gt;. It still hurts. And I still miss resources such as @TroveNewsBot and the Trove API Console which ran happily for more than a decade before being killed without warning by the NLA.&lt;/li&gt;&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html&#34;&gt;Saying goodbye to 17 years of work on the National Archives of Australia’s collections&lt;/a&gt;. This will be the first New Year’s Day in a decade when I haven’t updated my &lt;a href=&#34;https://updates.timsherratt.org/2025/02/05/ten-years-of-data-the.html&#34;&gt;harvest of files with the access status of closed&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;Next year&lt;/h2&gt;&lt;ul&gt;&lt;li&gt;In 2026, I’m looking forward to starting work on the &lt;a href=&#34;https://ardc.edu.au/project/reusable-and-accessible-public-interest-documents-rapid/&#34;&gt;RAPID project&lt;/a&gt;, building on the work I’ve done on Commonwealth Hansard over the years to create new examples and documentation.&lt;/li&gt;&lt;li&gt;I’m honoured to be giving the closing keynote at the &lt;a href=&#34;https://www.glamlabs.io/events/glam-labs-futures-26&#34;&gt;GLAM Labs Futures conference&lt;/a&gt; in Edinburgh in June — hoping we can pull together the funds to get there in person!&lt;/li&gt;&lt;/ul&gt;&lt;h2&gt;How you can help&lt;/h2&gt;&lt;p&gt;Much of my work is unfunded, and keeping resources such as the GLAM Name Index running costs real money. I’ve been very grateful for the support of my GitHub sponsors over past years. Their contributions help cover a substantial proportion of my cloud hosting costs. But bidding farewell to Twitter and Trove has had an impact on my sponsorship income. If you use or value the things I build to help researchers make use of GLAM collections, you might like to &lt;a href=&#34;https://github.com/sponsors/wragge&#34;&gt;sponsor me on GitHub&lt;/a&gt;, or &lt;a href=&#34;https://www.buymeacoffee.com/wragge&#34;&gt;Buy Me a Coffee&lt;/a&gt;. All contributions are greatly appreciated!&lt;/p&gt;&lt;p&gt;If you can’t afford a financial contribution, there are other ways you can help!&lt;/p&gt;&lt;ul&gt;&lt;li&gt;Let me know how you’re using my stuff! A bit of positive feedback does wonders when my enthusiasm is flagging. You can find my contact details at &lt;a href=&#34;https://timsherratt.au/&#34;&gt;timsherratt.au&lt;/a&gt;.&lt;/li&gt;&lt;li&gt;Tell others how you use my stuff! Getting information about resources out to those who might benefit is really hard, so your help would be greatly appreciated.&lt;/li&gt;&lt;li&gt;The GLAM Workbench describes a few other ways &lt;a href=&#34;https://glam-workbench.net/get-involved/&#34;&gt;you can get involved&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p&gt;Goodbye 2025!&lt;/p&gt;
</source:markdown>
    </item>
    
    <item>
      <title>Exploring Victorian newspapers</title>
      <link>https://updates.timsherratt.org/2025/12/16/exploring-victorian-newspapers.html</link>
      <pubDate>Tue, 16 Dec 2025 12:03:48 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/12/16/exploring-victorian-newspapers.html</guid>
      <description>&lt;p&gt;Newspapers are a vital source for local history. That&amp;rsquo;s why, &lt;a href=&#34;https://discontents.com.au/easter-eggsperiments/index.html&#34;&gt;back in 2014&lt;/a&gt;, I created the &lt;a href=&#34;https://wraggelabs.com/trove-places/map/&#34;&gt;Trove Places&lt;/a&gt; app – a map interface to help people find Trove&amp;rsquo;s digitised newspapers by their place of publication or distribution. Trove Places has proved very popular, and the State Libraries of South Australia, and Victoria, amongst others, point their users to it to help with their research. I&amp;rsquo;ve updated the data several times over the years, though the &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;Trove&amp;rsquo;s new gatekeeping regime&lt;/a&gt; will make future updates difficult.&lt;/p&gt;
&lt;p&gt;During &lt;a href=&#34;https://lab.slv.vic.gov.au/team/tim-sherratt&#34;&gt;my residency at the State Library of Victoria&lt;/a&gt;, one of the librarians noted how useful the app was, and asked whether it might be possible to include undigitised newspapers from the SLV catalogue as well as those in Trove. It was, and I did – here&amp;rsquo;s a &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;brand new app to explore Victorian newspapers&lt;/a&gt;, both digitised and undigitised!&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-12-12-12-48-16.png&#34; width=&#34;600&#34; height=&#34;391&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Just &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;click on the map&lt;/a&gt; to find Victorian newspapers!&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;It&amp;rsquo;s pretty easy to use. You just click on the map in an area you&amp;rsquo;re interested in. The map will display the 20 nearest places where newspapers where published or distributed. The size of the markers indicates how many titles are associated with each place. In the sidebar, details of the newspapers are listed by place, ordered by their distance from your selected point.&lt;/p&gt;
&lt;p&gt;You can also find local newspapers using the &lt;a href=&#34;https://slv.wraggelabs.com/myplace/&#34;&gt;my place&lt;/a&gt; app. Once you enter an address, newspapers from your suburb or town will be displayed, as well as those from nearby locations.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-12-16-12-51-54.png&#34; width=&#34;600&#34; height=&#34;507&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;newspapers from Geelong displayed in the my place app&lt;/figcaption&gt;&lt;/figure&gt;
&lt;h2 id=&#34;assembling-the-data&#34;&gt;Assembling the data&lt;/h2&gt;
&lt;p&gt;How do you find Victorian newspapers? The reference librarians at the SLV pointed me to the &amp;lsquo;Place newspaper published&amp;rsquo; field in the catalogue. Searching this field for &amp;lsquo;Australia&amp;ndash;Victoria&amp;rsquo; &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/search?query=lds03,exact,Australia--Victoria&amp;amp;tab=searchProfile&amp;amp;search_scope=slv_local&amp;amp;vid=61SLV_INST:SLV&#34;&gt;returns 3,997 results&lt;/a&gt;, compared to the 460 digitised in Trove.&lt;/p&gt;
&lt;p&gt;The first step in assembling the data was to harvest the newspaper records from the SLV catalogue. To do this I made use of the Primo JSON API. The method is &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_newspapers.ipynb&#34;&gt;documented in this notebook&lt;/a&gt;. The results was a &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/newspapers.ndjson&#34;&gt;newline-delimited JSON file&lt;/a&gt;, with one record per line.&lt;/p&gt;
&lt;p&gt;The harvested metadata doesn&amp;rsquo;t include links to digitised versions of newspapers in Trove. To add these links I first looked in the &lt;code&gt;856&lt;/code&gt; field of the newspaper&amp;rsquo;s MARC record. I also noticed that some Trove links were being loaded from an &amp;lsquo;edelivery&amp;rsquo; JSON file, so I added these as well. I ended up with 344 unique links to Trove, but not all of these were to digitised newspapers as some more recent newspapers are available through eLegal deposit. In total there were 268 unique links to digitised newspapers. This is well short of the 460 Victorian newspapers in Trove. Why? It&amp;rsquo;s possible that the links haven&amp;rsquo;t been added into the SLV catalogue, or that the &amp;lsquo;place newspaper published&amp;rsquo; field hasn&amp;rsquo;t been populated for records that include the links. It&amp;rsquo;s also possible that Trove links are hiding somewhere else in the SLV catalogue!&lt;/p&gt;
&lt;p&gt;To try and fill this gap, I compared the catalogue metadata with my &lt;a href=&#34;https://github.com/wragge/trove-newspaper-totals/blob/master/data/total_articles_by_newspaper.csv&#34;&gt;most recent harvest of Trove newspaper titles&lt;/a&gt;. If the Trove url was missing, I searched the catalogue data for the newspaper title. I then manually checked the results, making sure the dates and titles lined up, and added positive matches to &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/newspaper_manual_additions.csv&#34;&gt;a new CSV file&lt;/a&gt; which I merged back into the main dataset. This added another 152 Trove links.&lt;/p&gt;
&lt;p&gt;The next step was to link the &amp;lsquo;place newspaper published&amp;rsquo; values to places with known locations. The &amp;lsquo;place newspaper published&amp;rsquo; information is included in the &lt;code&gt;lds03&lt;/code&gt; field of the harvested metadata. Records often contain references to multiple places, so I split all the newspaper/place combinations out into separate rows. I then matched the places against a list of Victorian place names and coordinates downloaded from the &lt;a href=&#34;https://maps.land.vic.gov.au/lassi/VicnamesUI.jsp&#34;&gt;VicNames database&lt;/a&gt;. If there were no matches, I manually checked and adjusted the place names – for example, I changed &amp;lsquo;East Kew&amp;rsquo; to &amp;lsquo;Kew East&amp;rsquo;, and &amp;lsquo;Bayside&amp;rsquo; to &amp;lsquo;Bayside City&amp;rsquo;.&lt;/p&gt;
&lt;p&gt;To add any Trove digitised newspapers that might still be missing, I made use of my existing Trove harvests. First I compared my &lt;a href=&#34;https://docs.google.com/spreadsheets/d/1rURriHBSf3MocI8wsdl1114t0YeyU0BVSXWeg232MZs/edit?usp=sharing&#34;&gt;Trove Places dataset&lt;/a&gt; with my &lt;a href=&#34;https://github.com/wragge/trove-newspaper-totals/blob/master/data/total_articles_by_newspaper.csv&#34;&gt;latest harvest of newspaper titles&lt;/a&gt;. There were a few new titles, so I matched them to places &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/get_places_from_newspapers.ipynb&#34;&gt;using this notebook&lt;/a&gt;, based on my original Trove Places code. I then merged the Trove Places dataset with the new titles and checked it against the catalogue dataset. If any urls were missing, I added a record from the Trove data.&lt;/p&gt;
&lt;p&gt;All of the processing steps are &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_newspapers.ipynb&#34;&gt;documented in this notebook&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id=&#34;building-the-apps&#34;&gt;Building the apps&lt;/h2&gt;
&lt;p&gt;To make the data easily searchable by its geospatial coordinates, I loaded all the data into an SQLite/Spatialite database and &lt;a href=&#34;https://slv-places-481615284700.australia-southeast1.run.app/newspapers&#34;&gt;published it online using Datasette&lt;/a&gt;. The database contains linked tables for titles and places.&lt;/p&gt;
&lt;p&gt;I also created a couple of canned queries which, together with Datasette&amp;rsquo;s built-in JSON API, made it possible to retrieve places and titles based on their distance from a given point. For example, this url retrieves places ordered by their distance from the point at latitude -36.815, longitude 144.965 : &lt;a href=&#34;https://slv-places-481615284700.australia-southeast1.run.app/newspapers/places_from_point.json?longitude=144.965&amp;amp;latitude=-36.815&amp;amp;distance=100000&amp;amp;_shape=array&#34;&gt;https://slv-places-481615284700.australia-southeast1.run.app/newspapers/places_from_point.json?longitude=144.965&amp;amp;latitude=-36.815&amp;amp;distance=100000&amp;amp;_shape=array&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;When you click on the map in the &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;Victorian Newspapers Explorer&lt;/a&gt;, it fires off a request like this to find nearby places. It then makes a second request to find newspapers related to those places and displays the results.&lt;/p&gt;
&lt;p&gt;The Victorian Newspapers Explorer was my first attempt at using &lt;a href=&#34;https://maplibre.org&#34;&gt;MapLibre&lt;/a&gt; rather than Leaflet to display maps using Javascript. It&amp;rsquo;s more verbose, but more flexible, so I think I&amp;rsquo;ll gradually switch over my other apps, including Trove Places.&lt;/p&gt;
&lt;p&gt;All the code of the Victorian Newspapers Explorer is in the &lt;a href=&#34;https://github.com/wragge/slv-demo-apps&#34;&gt;slv-demo-apps repository&lt;/a&gt;.&lt;/p&gt;
</description>
      <source:markdown>

Newspapers are a vital source for local history. That&#39;s why, [back in 2014](https://discontents.com.au/easter-eggsperiments/index.html), I created the [Trove Places](https://wraggelabs.com/trove-places/map/) app – a map interface to help people find Trove&#39;s digitised newspapers by their place of publication or distribution. Trove Places has proved very popular, and the State Libraries of South Australia, and Victoria, amongst others, point their users to it to help with their research. I&#39;ve updated the data several times over the years, though the [Trove&#39;s new gatekeeping regime](https://updates.timsherratt.org/2025/05/07/farewell-trove.html) will make future updates difficult.

During [my residency at the State Library of Victoria](https://lab.slv.vic.gov.au/team/tim-sherratt), one of the librarians noted how useful the app was, and asked whether it might be possible to include undigitised newspapers from the SLV catalogue as well as those in Trove. It was, and I did – here&#39;s a [brand new app to explore Victorian newspapers](https://slv.wraggelabs.com/newspapers/), both digitised and undigitised!

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-12-12-12-48-16.png&#34; width=&#34;600&#34; height=&#34;391&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Just &lt;a href=&#34;https://slv.wraggelabs.com/newspapers/&#34;&gt;click on the map&lt;/a&gt; to find Victorian newspapers!&lt;/figcaption&gt;&lt;/figure&gt;

It&#39;s pretty easy to use. You just click on the map in an area you&#39;re interested in. The map will display the 20 nearest places where newspapers where published or distributed. The size of the markers indicates how many titles are associated with each place. In the sidebar, details of the newspapers are listed by place, ordered by their distance from your selected point.

You can also find local newspapers using the [my place](https://slv.wraggelabs.com/myplace/) app. Once you enter an address, newspapers from your suburb or town will be displayed, as well as those from nearby locations. 

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-12-16-12-51-54.png&#34; width=&#34;600&#34; height=&#34;507&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;newspapers from Geelong displayed in the my place app&lt;/figcaption&gt;&lt;/figure&gt;

## Assembling the data

How do you find Victorian newspapers? The reference librarians at the SLV pointed me to the &#39;Place newspaper published&#39; field in the catalogue. Searching this field for &#39;Australia--Victoria&#39; [returns 3,997 results](https://find.slv.vic.gov.au/discovery/search?query=lds03,exact,Australia--Victoria&amp;tab=searchProfile&amp;search_scope=slv_local&amp;vid=61SLV_INST:SLV), compared to the 460 digitised in Trove.

The first step in assembling the data was to harvest the newspaper records from the SLV catalogue. To do this I made use of the Primo JSON API. The method is [documented in this notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_newspapers.ipynb). The results was a [newline-delimited JSON file](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/newspapers.ndjson), with one record per line.

The harvested metadata doesn&#39;t include links to digitised versions of newspapers in Trove. To add these links I first looked in the `856` field of the newspaper&#39;s MARC record. I also noticed that some Trove links were being loaded from an &#39;edelivery&#39; JSON file, so I added these as well. I ended up with 344 unique links to Trove, but not all of these were to digitised newspapers as some more recent newspapers are available through eLegal deposit. In total there were 268 unique links to digitised newspapers. This is well short of the 460 Victorian newspapers in Trove. Why? It&#39;s possible that the links haven&#39;t been added into the SLV catalogue, or that the &#39;place newspaper published&#39; field hasn&#39;t been populated for records that include the links. It&#39;s also possible that Trove links are hiding somewhere else in the SLV catalogue!

To try and fill this gap, I compared the catalogue metadata with my [most recent harvest of Trove newspaper titles](https://github.com/wragge/trove-newspaper-totals/blob/master/data/total_articles_by_newspaper.csv). If the Trove url was missing, I searched the catalogue data for the newspaper title. I then manually checked the results, making sure the dates and titles lined up, and added positive matches to [a new CSV file](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/newspaper_manual_additions.csv) which I merged back into the main dataset. This added another 152 Trove links.

The next step was to link the &#39;place newspaper published&#39; values to places with known locations. The &#39;place newspaper published&#39; information is included in the `lds03` field of the harvested metadata. Records often contain references to multiple places, so I split all the newspaper/place combinations out into separate rows. I then matched the places against a list of Victorian place names and coordinates downloaded from the [VicNames database](https://maps.land.vic.gov.au/lassi/VicnamesUI.jsp). If there were no matches, I manually checked and adjusted the place names – for example, I changed &#39;East Kew&#39; to &#39;Kew East&#39;, and &#39;Bayside&#39; to &#39;Bayside City&#39;.

To add any Trove digitised newspapers that might still be missing, I made use of my existing Trove harvests. First I compared my [Trove Places dataset](https://docs.google.com/spreadsheets/d/1rURriHBSf3MocI8wsdl1114t0YeyU0BVSXWeg232MZs/edit?usp=sharing) with my [latest harvest of newspaper titles](https://github.com/wragge/trove-newspaper-totals/blob/master/data/total_articles_by_newspaper.csv). There were a few new titles, so I matched them to places [using this notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/get_places_from_newspapers.ipynb), based on my original Trove Places code. I then merged the Trove Places dataset with the new titles and checked it against the catalogue dataset. If any urls were missing, I added a record from the Trove data.

All of the processing steps are [documented in this notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_newspapers.ipynb).

## Building the apps

To make the data easily searchable by its geospatial coordinates, I loaded all the data into an SQLite/Spatialite database and [published it online using Datasette](https://slv-places-481615284700.australia-southeast1.run.app/newspapers). The database contains linked tables for titles and places.

I also created a couple of canned queries which, together with Datasette&#39;s built-in JSON API, made it possible to retrieve places and titles based on their distance from a given point. For example, this url retrieves places ordered by their distance from the point at latitude -36.815, longitude 144.965 : https://slv-places-481615284700.australia-southeast1.run.app/newspapers/places_from_point.json?longitude=144.965&amp;latitude=-36.815&amp;distance=100000&amp;_shape=array 

When you click on the map in the [Victorian Newspapers Explorer](https://slv.wraggelabs.com/newspapers/), it fires off a request like this to find nearby places. It then makes a second request to find newspapers related to those places and displays the results.

The Victorian Newspapers Explorer was my first attempt at using [MapLibre](https://maplibre.org) rather than Leaflet to display maps using Javascript. It&#39;s more verbose, but more flexible, so I think I&#39;ll gradually switch over my other apps, including Trove Places.

All the code of the Victorian Newspapers Explorer is in the [slv-demo-apps repository](https://github.com/wragge/slv-demo-apps).




</source:markdown>
    </item>
    
    <item>
      <title>Why bother?</title>
      <link>https://updates.timsherratt.org/2025/12/03/why-bother.html</link>
      <pubDate>Wed, 03 Dec 2025 15:18:22 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/12/03/why-bother.html</guid>
      <description>&lt;p&gt;&lt;em&gt;This was the introduction to my talk on the results of my time as Creative Technologist-in-Residence at the State Library of Victoria. My slides, with my full notes &lt;a href=&#34;https://slides.com/wragge/slv-my-place&#34;&gt;are available online&lt;/a&gt;, but after a very strange year that has travelled from disappointment to exhilaration, I thought it was worth posting these words separately.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;The work that I do, that I&amp;rsquo;ve been doing for the past 30 years, is focused on helping people find, use, and understand the wonderfully rich collections held by our libraries, archives, and museums – the GLAM sector. Much of it is quite practical, resulting in tools and applications that are used by a wide range of researchers. Some of it is playful, some of it is critical, and some of it is just weird.&lt;/p&gt;
&lt;p&gt;You can browse through some of this history on &lt;a href=&#34;https://wraggelabs.com&#34;&gt;wraggelabs.com&lt;/a&gt;. And if you&amp;rsquo;re interested in my current crop of tools you can head to the &lt;a href=&#34;https://glam-workbench.net&#34;&gt;GLAM Workbench&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;As some of you may know, I had &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;a few setbacks&lt;/a&gt; at the beginning of this year which really made me wonder whether I wanted to continue doing this sort of work.&lt;/p&gt;
&lt;p&gt;I mean, why bother?&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m really grateful that &lt;a href=&#34;https://slv.wraggelabs.com&#34;&gt;this residency&lt;/a&gt; has given me a chance to refocus on the reasons why I do what I do.&lt;/p&gt;
&lt;p&gt;I suppose my starting point is the fact that libraries can&amp;rsquo;t do everything themselves.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m thinking here specifically about the digital research space. There&amp;rsquo;s a lot that libraries, and other GLAM organisations, &lt;em&gt;can&lt;/em&gt; do – provide search interfaces, APIs, downloadable datasets, documentation, and examples of how to access APIs and datasets using code. The sorts of things that the &lt;a href=&#34;https://lab.slv.vic.gov.au&#34;&gt;SLV LAB&lt;/a&gt; is doing.&lt;/p&gt;
&lt;p&gt;I should pause here to unpack some acronyms. APIs deliver data in a form that machines can understand and process. Websites are for humans, APIs are for computers. APIs are also building blocks which can be connected up to create new applications – and I&amp;rsquo;ll be showing some examples of this later on.&lt;/p&gt;
&lt;p&gt;So there is much that GLAM organisations can do to support digital research. But it will never be enough. Researchers – whether they be academics or family historians – will always want more. It is in the nature of research to ask new questions, to head off in new directions.&lt;/p&gt;
&lt;p&gt;But rather than see this as a source of tension, I see it as an opportunity for collaboration. An opportunity to cultivate the &lt;em&gt;in-between&lt;/em&gt; spaces where research methods, tools, and results can feed back into the contextual frameworks of GLAM collections. Where GLAM organisations can share and celebrate the work that&amp;rsquo;s done with their data. Where all can find inspiration, ideas and support.&lt;/p&gt;
&lt;p&gt;We&amp;rsquo;ve tended to call this sort of stuff infrastructure, but I think that really downplays the human aspect. The research sector has started to develop the funding and career structures necessary to allow people to build and maintain these infrastructures, but we need more. We need to recognise that a single tool, developed by an individual without institutional support, can be just as important as a multi-million dollar platform. Passion is precious and needs to be protected.&lt;/p&gt;
&lt;p&gt;Most of all, we need to keep a focus on the ethical imperatives – the reasons &lt;em&gt;why&lt;/em&gt; we bother and &lt;em&gt;why&lt;/em&gt; it matters. For me it boils down to openness and generosity. I have benefited greatly from the openness and generosity of others, and I want to pass that on. It&amp;rsquo;s the glue we need to hold those in-between spaces together; the sustenance we need to maintain our enthusiasm in the face of all the crap; the inspiration we need to try something new.&lt;/p&gt;
&lt;p&gt;Initiatives like the SLV LAB are important, not just because they foster innovation, but because they invite new ideas in. They even give space for ageing hackers like me to spend some dedicated time doing what they love – crafting new pathways for people to explore our glorious GLAM collections.&lt;/p&gt;
</description>
      <source:markdown>*This was the introduction to my talk on the results of my time as Creative Technologist-in-Residence at the State Library of Victoria. My slides, with my full notes [are available online](https://slides.com/wragge/slv-my-place), but after a very strange year that has travelled from disappointment to exhilaration, I thought it was worth posting these words separately.*

----

The work that I do, that I&#39;ve been doing for the past 30 years, is focused on helping people find, use, and understand the wonderfully rich collections held by our libraries, archives, and museums – the GLAM sector. Much of it is quite practical, resulting in tools and applications that are used by a wide range of researchers. Some of it is playful, some of it is critical, and some of it is just weird.

You can browse through some of this history on [wraggelabs.com](https://wraggelabs.com). And if you&#39;re interested in my current crop of tools you can head to the [GLAM Workbench](https://glam-workbench.net).

As some of you may know, I had [a few setbacks](https://updates.timsherratt.org/2025/05/07/farewell-trove.html) at the beginning of this year which really made me wonder whether I wanted to continue doing this sort of work.

I mean, why bother?

I&#39;m really grateful that [this residency](https://slv.wraggelabs.com) has given me a chance to refocus on the reasons why I do what I do.

I suppose my starting point is the fact that libraries can&#39;t do everything themselves. 

I&#39;m thinking here specifically about the digital research space. There&#39;s a lot that libraries, and other GLAM organisations, *can* do – provide search interfaces, APIs, downloadable datasets, documentation, and examples of how to access APIs and datasets using code. The sorts of things that the [SLV LAB](https://lab.slv.vic.gov.au) is doing.

I should pause here to unpack some acronyms. APIs deliver data in a form that machines can understand and process. Websites are for humans, APIs are for computers. APIs are also building blocks which can be connected up to create new applications – and I&#39;ll be showing some examples of this later on.

So there is much that GLAM organisations can do to support digital research. But it will never be enough. Researchers – whether they be academics or family historians – will always want more. It is in the nature of research to ask new questions, to head off in new directions.

But rather than see this as a source of tension, I see it as an opportunity for collaboration. An opportunity to cultivate the *in-between* spaces where research methods, tools, and results can feed back into the contextual frameworks of GLAM collections. Where GLAM organisations can share and celebrate the work that&#39;s done with their data. Where all can find inspiration, ideas and support.

We&#39;ve tended to call this sort of stuff infrastructure, but I think that really downplays the human aspect. The research sector has started to develop the funding and career structures necessary to allow people to build and maintain these infrastructures, but we need more. We need to recognise that a single tool, developed by an individual without institutional support, can be just as important as a multi-million dollar platform. Passion is precious and needs to be protected.

Most of all, we need to keep a focus on the ethical imperatives – the reasons *why* we bother and *why* it matters. For me it boils down to openness and generosity. I have benefited greatly from the openness and generosity of others, and I want to pass that on. It&#39;s the glue we need to hold those in-between spaces together; the sustenance we need to maintain our enthusiasm in the face of all the crap; the inspiration we need to try something new.

Initiatives like the SLV LAB are important, not just because they foster innovation, but because they invite new ideas in. They even give space for ageing hackers like me to spend some dedicated time doing what they love – crafting new pathways for people to explore our glorious GLAM collections. 

</source:markdown>
    </item>
    
    <item>
      <title>Counting down... (to the end of my SLV residency)</title>
      <link>https://updates.timsherratt.org/2025/11/19/counting-down-to-the-end.html</link>
      <pubDate>Wed, 19 Nov 2025 15:01:22 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/11/19/counting-down-to-the-end.html</guid>
      <description>&lt;p&gt;My stint as &lt;a href=&#34;https://updates.timsherratt.org/2025/09/22/creative-technologistinresidence-at-the-state.html&#34;&gt;Creative Technologist-in-Residence at the State Library of Victoria LAB&lt;/a&gt; comes to an end in a few weeks time and I&amp;rsquo;m frantically trying to pull things together. I&amp;rsquo;ll be back on-site at the Library from 1 to 5 December for a few events, and to report back to staff on what I&amp;rsquo;ve been doing.&lt;/p&gt;
&lt;p&gt;On Tuesday &lt;strong&gt;2 December&lt;/strong&gt;, there&amp;rsquo;ll be a public workshop on using and contributing to the &lt;a href=&#34;https://glam-workbench.net&#34;&gt;GLAM Workbench&lt;/a&gt;. Here&amp;rsquo;s the blurb:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;More and more GLAM organisations are looking to share their data to foster creativity and support new types of research. But how can you help potential users understand the possibilities of your data? This workshop will explore how GLAM organisations can create and share resources that encourage experimentation.&lt;/p&gt;
&lt;p&gt;The GLAM Workbench is a large collection of tools, hacks, and tutorials aimed at helping researchers make use of collection data. It uses platforms such as Jupyter notebooks to create live, working examples that run in your browser without additional software. Similar repositories of computational resources are being developed by GLAM organisations around the world.&lt;/p&gt;
&lt;p&gt;This workshop will introduce the technologies and standards used in the GLAM Workbench, such as Jupyter notebooks. It will provide an overview of related activity around the world, including best practice guidelines for GLAM organisations developing computational resources. It will explain how organisations and individuals can contribute content to the GLAM Workbench, or use it as a model to create their own specialised workbenches.&lt;/p&gt;
&lt;p&gt;Sharing data is important, but so is sharing skills, tools, and knowledge. Come along to find out how the GLAM Workbench can help.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;It&amp;rsquo;s a free, hybrid event (in person and online) and will run from 1.00-3.00pm. A sign up page should be available soon.&lt;/p&gt;
&lt;p&gt;On Wednesday &lt;strong&gt;3 December&lt;/strong&gt; I&amp;rsquo;m presenting the results of my residency in a &amp;lsquo;technologist&amp;rsquo;s talk&amp;rsquo;. It&amp;rsquo;s an internal event, but it&amp;rsquo;s in the public &amp;lsquo;Create quarter&amp;rsquo; of the Library, so I think anyone can pop in. Hopefully there&amp;rsquo;ll be a video I can share.&lt;/p&gt;
&lt;p&gt;To give you an idea of what I&amp;rsquo;ll be talking about, here&amp;rsquo;s some of the outcomes so far:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;hacking the library workshop (&lt;a href=&#34;https://slides.com/wragge/slv-code-club&#34;&gt;slides&lt;/a&gt;, and &lt;a href=&#34;https://updates.timsherratt.org/2025/09/23/exploring-slv-urls.html&#34;&gt;blog post about urls&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;bounding boxes for parish maps (&lt;a href=&#34;https://updates.timsherratt.org/2025/10/06/creating-bounding-boxes-for-parish.html&#34;&gt;blog post&lt;/a&gt;, &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency&#34;&gt;code&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;geolocating the Committee for Urban Action collection of photographs (&lt;a href=&#34;https://wragge.github.io/slv-demo-apps/cua-browser.html&#34;&gt;prototype interface&lt;/a&gt;, still documenting the method)&lt;/li&gt;
&lt;li&gt;a new fully-searchable version of the Sands &amp;amp; MacDougall&amp;rsquo;s directories (&lt;a href=&#34;https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html&#34;&gt;blog post&lt;/a&gt;, &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;database&lt;/a&gt;, and &lt;a href=&#34;https://updates.timsherratt.org/2025/11/16/some-sands-mac-tweaks-thanks.html&#34;&gt;another blog post&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;georeferencing digitised maps – over 500 so far! (&lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;documentation&lt;/a&gt;, &lt;a href=&#34;https://wragge.github.io/slv-allmaps/dashboard.html&#34;&gt;dashboard&lt;/a&gt;, &lt;a href=&#34;https://github.com/wragge/slv-allmaps&#34;&gt;data repository&lt;/a&gt;, &lt;a href=&#34;https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html&#34;&gt;blog post&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;and as of yesterday, 3,000+ geolocated newspapers (documentation and interface coming!)&lt;/li&gt;
&lt;/ul&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-18-15-22-26.png&#34; width=&#34;600&#34; height=&#34;356&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;First attempt at mapping places of publication and distribution of Victorian newspapers&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;At the moment I&amp;rsquo;m trying to bring it all together in a new interface that let&amp;rsquo;s you type in an address and find collection materials relating to your home, your street, and your suburb. Only two weeks to go! Eeek!&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-15-17-30-26.png&#34; width=&#34;600&#34; height=&#34;454&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Work in progress!&lt;/figcaption&gt;&lt;/figure&gt;
</description>
      <source:markdown>My stint as [Creative Technologist-in-Residence at the State Library of Victoria LAB](https://updates.timsherratt.org/2025/09/22/creative-technologistinresidence-at-the-state.html) comes to an end in a few weeks time and I&#39;m frantically trying to pull things together. I&#39;ll be back on-site at the Library from 1 to 5 December for a few events, and to report back to staff on what I&#39;ve been doing.

On Tuesday **2 December**, there&#39;ll be a public workshop on using and contributing to the [GLAM Workbench](https://glam-workbench.net). Here&#39;s the blurb:

&gt; More and more GLAM organisations are looking to share their data to foster creativity and support new types of research. But how can you help potential users understand the possibilities of your data? This workshop will explore how GLAM organisations can create and share resources that encourage experimentation.
&gt;
&gt; The GLAM Workbench is a large collection of tools, hacks, and tutorials aimed at helping researchers make use of collection data. It uses platforms such as Jupyter notebooks to create live, working examples that run in your browser without additional software. Similar repositories of computational resources are being developed by GLAM organisations around the world.
&gt;
&gt; This workshop will introduce the technologies and standards used in the GLAM Workbench, such as Jupyter notebooks. It will provide an overview of related activity around the world, including best practice guidelines for GLAM organisations developing computational resources. It will explain how organisations and individuals can contribute content to the GLAM Workbench, or use it as a model to create their own specialised workbenches.
&gt;
&gt; Sharing data is important, but so is sharing skills, tools, and knowledge. Come along to find out how the GLAM Workbench can help.

It&#39;s a free, hybrid event (in person and online) and will run from 1.00-3.00pm. A sign up page should be available soon.

On Wednesday **3 December** I&#39;m presenting the results of my residency in a &#39;technologist&#39;s talk&#39;. It&#39;s an internal event, but it&#39;s in the public &#39;Create quarter&#39; of the Library, so I think anyone can pop in. Hopefully there&#39;ll be a video I can share.

To give you an idea of what I&#39;ll be talking about, here&#39;s some of the outcomes so far:

- hacking the library workshop ([slides](https://slides.com/wragge/slv-code-club), and [blog post about urls](https://updates.timsherratt.org/2025/09/23/exploring-slv-urls.html))
- bounding boxes for parish maps ([blog post](https://updates.timsherratt.org/2025/10/06/creating-bounding-boxes-for-parish.html), [code](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency))
- geolocating the Committee for Urban Action collection of photographs ([prototype interface](https://wragge.github.io/slv-demo-apps/cua-browser.html), still documenting the method)
- a new fully-searchable version of the Sands &amp; MacDougall&#39;s directories ([blog post](https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html), [database](https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/), and [another blog post](https://updates.timsherratt.org/2025/11/16/some-sands-mac-tweaks-thanks.html))
- georeferencing digitised maps – over 500 so far! ([documentation](https://wragge.github.io/slv-allmaps/), [dashboard](https://wragge.github.io/slv-allmaps/dashboard.html), [data repository](https://github.com/wragge/slv-allmaps), [blog post](https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html))
- and as of yesterday, 3,000+ geolocated newspapers (documentation and interface coming!)

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-18-15-22-26.png&#34; width=&#34;600&#34; height=&#34;356&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;First attempt at mapping places of publication and distribution of Victorian newspapers&lt;/figcaption&gt;&lt;/figure&gt;

At the moment I&#39;m trying to bring it all together in a new interface that let&#39;s you type in an address and find collection materials relating to your home, your street, and your suburb. Only two weeks to go! Eeek!

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-15-17-30-26.png&#34; width=&#34;600&#34; height=&#34;454&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Work in progress!&lt;/figcaption&gt;&lt;/figure&gt;




</source:markdown>
    </item>
    
    <item>
      <title>Some Sands &amp; Mac tweaks thanks to ALTO and IIIF</title>
      <link>https://updates.timsherratt.org/2025/11/16/some-sands-mac-tweaks-thanks.html</link>
      <pubDate>Sun, 16 Nov 2025 14:17:26 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/11/16/some-sands-mac-tweaks-thanks.html</guid>
      <description>&lt;p&gt;I posted recently about &lt;a href=&#34;https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html&#34;&gt;my new fully-searchable version of the Sands &amp;amp; MacDougall directories&lt;/a&gt;. I&amp;rsquo;ve now moved on to try and pull together a number of the State Library of Victoria&amp;rsquo;s place-based collections into a new discovery interface. It&amp;rsquo;s going to be a busy couple of weeks as my residency ends in early December!&lt;/p&gt;
&lt;p&gt;I wanted to incorporate Sands &amp;amp; Mac search results into the new interface. Getting the data was easy because &lt;a href=&#34;https://datasette.io&#34;&gt;Datasette&lt;/a&gt; has a JSON API baked in. But what about the images? I could just display a thumbnail of the whole page, but it would be better to show a snippet of the actual entry. Thanks to &lt;a href=&#34;https://iiif.io&#34;&gt;IIIF&lt;/a&gt; and &lt;a href=&#34;https://en.wikipedia.org/wiki/Analyzed_Layout_and_Text_Object&#34;&gt;ALTO&lt;/a&gt;, I now can.&lt;/p&gt;
&lt;p&gt;IIIF makes it easy to cut small sections out of a larger image. You just put the coordinates of the desired section in the IIIF url. As I noted in my previous post, the ALTO files that contain the OCR data from Sands &amp;amp; Mac include the coordinates of every line, and every word. I just had to bring the two together.&lt;/p&gt;
&lt;p&gt;All I did was update the code that extracts the data from the ALTO files to save the results as newline delimited JSON instead of a plain text file. Each line in each volume of Sands &amp;amp; Mac is now saved a JSON object that contains the text, as well as the height, width, vertical position, and horizontal position of the line within the page image. When I load up the SQLite database, I add the values for &lt;code&gt;h&lt;/code&gt;, &lt;code&gt;w&lt;/code&gt;, &lt;code&gt;x&lt;/code&gt;, and &lt;code&gt;y&lt;/code&gt; as well as the text for each line.&lt;/p&gt;
&lt;p&gt;What does this make possible?&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;When you go to an individual entry, the page image now automatically pans and zooms so that the current entry is at the centre of the image viewer. I just updated the &lt;a href=&#34;https://openseadragon.github.io&#34;&gt;OpenSeadragon&lt;/a&gt; code to focus on the entry&amp;rsquo;s position.&lt;/li&gt;
&lt;li&gt;If you share an entry on social media, a snipped out section of the page image showing the selected entry is displayed as there&amp;rsquo;s now an image &lt;code&gt;META&lt;/code&gt; tag that points to an IIIF url.&lt;/li&gt;
&lt;li&gt;You can retrieve entries via the API and use the coordinates to request snipped out images of them via IIIF.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-16-15-06-52.png&#34; width=&#34;600&#34; height=&#34;330&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Nice image snippets thanks to IIIF and ALTO (and a sneak prview of what&#39;s coming...)&lt;/figcaption&gt;&lt;/figure&gt;
</description>
      <source:markdown>I posted recently about [my new fully-searchable version of the Sands &amp; MacDougall directories](https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html). I&#39;ve now moved on to try and pull together a number of the State Library of Victoria&#39;s place-based collections into a new discovery interface. It&#39;s going to be a busy couple of weeks as my residency ends in early December!

I wanted to incorporate Sands &amp; Mac search results into the new interface. Getting the data was easy because [Datasette](https://datasette.io) has a JSON API baked in. But what about the images? I could just display a thumbnail of the whole page, but it would be better to show a snippet of the actual entry. Thanks to [IIIF](https://iiif.io) and [ALTO](https://en.wikipedia.org/wiki/Analyzed_Layout_and_Text_Object), I now can.

IIIF makes it easy to cut small sections out of a larger image. You just put the coordinates of the desired section in the IIIF url. As I noted in my previous post, the ALTO files that contain the OCR data from Sands &amp; Mac include the coordinates of every line, and every word. I just had to bring the two together.

All I did was update the code that extracts the data from the ALTO files to save the results as newline delimited JSON instead of a plain text file. Each line in each volume of Sands &amp; Mac is now saved a JSON object that contains the text, as well as the height, width, vertical position, and horizontal position of the line within the page image. When I load up the SQLite database, I add the values for `h`, `w`, `x`, and `y` as well as the text for each line.

What does this make possible?

1. When you go to an individual entry, the page image now automatically pans and zooms so that the current entry is at the centre of the image viewer. I just updated the [OpenSeadragon](https://openseadragon.github.io) code to focus on the entry&#39;s position.
2. If you share an entry on social media, a snipped out section of the page image showing the selected entry is displayed as there&#39;s now an image `META` tag that points to an IIIF url.
3. You can retrieve entries via the API and use the coordinates to request snipped out images of them via IIIF.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-11-16-15-06-52.png&#34; width=&#34;600&#34; height=&#34;330&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Nice image snippets thanks to IIIF and ALTO (and a sneak prview of what&#39;s coming...)&lt;/figcaption&gt;&lt;/figure&gt;


</source:markdown>
    </item>
    
    <item>
      <title>A new way of searching Sands &amp; Mac</title>
      <link>https://updates.timsherratt.org/2025/11/12/a-new-way-of-searching.html</link>
      <pubDate>Wed, 12 Nov 2025 21:25:21 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/11/12/a-new-way-of-searching.html</guid>
      <description>&lt;p&gt;In the fortnight I spent onsite at the State Library of Victoria, &amp;lsquo;Sands &amp;amp; Mac&amp;rsquo; was mentioned many times. And no wonder. The &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;amp;collectionId=81213035910007636&#34;&gt;Sands &amp;amp; McDougall&amp;rsquo;s directories&lt;/a&gt; are a goldmine for anyone researching family, local, or social history. They list thousands of names and addresses, enabling you to find individuals, and explore changing land use over time. When people ask the SLV&amp;rsquo;s librarians, &amp;lsquo;What can you tell me about the history of my house?&amp;rsquo;, Sands &amp;amp; Mac is one of the first resources consulted.&lt;/p&gt;
&lt;p&gt;The SLV has digitised &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;amp;collectionId=81213035910007636&#34;&gt;24 volumes of Sands &amp;amp; Mac&lt;/a&gt;, one every five years from 1860 to 1974. You can browse the contents of each volume in the SLV image viewer, using the partial contents listing to help you find your way to sections of interest. To search the full text content you need to use the PDF version, either in the built-in viewer, or by downloading the PDF. There&amp;rsquo;s a &lt;a href=&#34;https://blogs.slv.vic.gov.au/tips-and-tricks/collection-discovery-tips-sands-mcdougalls-directories/&#34;&gt;handy guide to using Sands &amp;amp; Mac&lt;/a&gt; that explains the options.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;However, there&amp;rsquo;s currently no way of searching across all 24 volumes, so as part of my residency at the SLV LAB, I thought I&amp;rsquo;d make one!&lt;/strong&gt;&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac.png&#34; width=&#34;600&#34; height=&#34;310&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;&lt;b&gt;Try it now!&lt;/b&gt;&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;My &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;new Sands &amp;amp; Mac database&lt;/a&gt; follows the pattern I&amp;rsquo;ve used previously to create fully-searchable versions of the &lt;a href=&#34;https://glam-workbench.net/trove-journals/nsw-post-office-directories/&#34;&gt;NSW Post Office directories&lt;/a&gt;, &lt;a href=&#34;https://glam-workbench.net/trove-journals/sydney-telephone-directories/&#34;&gt;Sydney telephone directories&lt;/a&gt;, and &lt;a href=&#34;https://glam-workbench.net/tasmanian-post-office-directories/&#34;&gt;Tasmanian Post Office directories&lt;/a&gt;. Every line of text is saved to a database, so a single query searches for entries across all volumes. You can also use advanced search features like wildcards and boolean operators.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac-search.png&#34; width=&#34;600&#34; height=&#34;543&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Search across all 24 volumes!&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;Once you&amp;rsquo;ve found a relevant entry you can view it in context, alongside a zoomable image of the page. You can even use Zotero to save individual entries to your own research database. &lt;a href=&#34;https://chineseaustralia.org/from-the-archive-uncovering-the-everyday-heritage-of-chinese-tasmanians/&#34;&gt;This blog post&lt;/a&gt; from the Everyday Heritage project describes how the Tasmanian directories have been used to map Tasmania&amp;rsquo;s Chinese population.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac-entry.png&#34; width=&#34;600&#34; height=&#34;370&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;View each entry in context! (Here&#39;s my Dad building his first house in Beaumaris in the 1950s.)&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;There&amp;rsquo;s still a few things I&amp;rsquo;d like to try, such as making use of the table of contents information for each volume. I&amp;rsquo;d also like to create some additional entry points to take users directly to listings for individual suburbs (maybe even streets!). Each volume has a directory of suburbs, so it would be a matter of extracting and cleaning the data and linking the entries to digitised pages. Certainly possible, but I don&amp;rsquo;t think I&amp;rsquo;ll have time to get it all done before the end of my residency. Perhaps I&amp;rsquo;ll try to get at least one volume done to demonstrate how it might work, and the value it would add. As I was writing this blog post I also realised there&amp;rsquo;s &lt;a href=&#34;https://www.environment.vic.gov.au/sustainability/victoria-unearthed/about-the-data/sands-and-mcdougall&#34;&gt;a dataset of businesses&lt;/a&gt; extracted from the Sands &amp;amp; Mac, so I need to think about how I can use that as well!&lt;/p&gt;
&lt;h2 id=&#34;technical-information-follows&#34;&gt;Technical information follows&amp;hellip;&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;ve documented the process I used to create fully-searchable versions of the &lt;a href=&#34;https://glam-workbench.net/libraries-tasmania/&#34;&gt;Tasmanian&lt;/a&gt; and &lt;a href=&#34;https://glam-workbench.net/trove-journals/create-text-db-indexed-by-line/&#34;&gt;NSW directories&lt;/a&gt;  in the GLAM Workbench. I followed a similar method for Sands and Mac, though with a few dead-ends and discoveries along the way.&lt;/p&gt;
&lt;h3 id=&#34;downloading-the-pdfs&#34;&gt;Downloading the PDFs&lt;/h3&gt;
&lt;p&gt;I assumed that it would be easiest to work from the PDF versions of each volume, as I&amp;rsquo;d done for Tasmania. So I set about finding a way to download them all. There&amp;rsquo;s only 24 volumes, so I &lt;em&gt;could&lt;/em&gt; have downloaded them manually, but where&amp;rsquo;s the fun in that?&lt;/p&gt;
&lt;p&gt;I started with a CSV file listing the Sands &amp;amp; Mac volumes that I downloaded from the catalogue. This gave me the Alma identifiers for each volume. To download the PDFs I needed two more identifiers, the &lt;code&gt;IE&lt;/code&gt; identifier assigned to each digitised item, and a file identifier that points to the PDF version of the item. The &lt;code&gt;IE&lt;/code&gt; identifier can be extracted from the item&amp;rsquo;s MARC record, as I described in &lt;a href=&#34;https://updates.timsherratt.org/2025/09/23/exploring-slv-urls.html&#34;&gt;my post on exploring urls&lt;/a&gt;. The PDF file identifier was a bit more difficult to track down. The PDF links in the image viewer are generated dynamically, so the data had to be coming from somewhere. Eventually I found that the viewer loaded a JSON file with all sorts of useful metadata in it!&lt;/p&gt;
&lt;p&gt;The url to download the JSON file is: &lt;code&gt;https://viewerapi.slv.vic.gov.au/?entity=[IE identifier]&amp;amp;dc_arrays=1&lt;/code&gt;. In the &lt;code&gt;summary&lt;/code&gt; section I found identifiers for &lt;code&gt;small_pdf&lt;/code&gt; and &lt;code&gt;master_pdf&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;I could then use these identifiers to construct urls to download the PDFs themselves: &lt;code&gt;https://rosetta.slv.vic.gov.au/delivery/DeliveryManagerServlet?dps_func=stream&amp;amp;dps_pid=[PDF id]&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;Once I had the PDFs I used &lt;a href=&#34;https://github.com/pymupdf/PyMuPDF&#34;&gt;PyMuPDF&lt;/a&gt; to extract all the text and images. As I suspected the text wasn&amp;rsquo;t really fit for purpose. The OCR was ok, but the column structures were a mess. Because I wanted to index each entry individually, it was important to try and get the columns represented as accurately as possible. The images in the small PDFs were already bitonal, so I started feeding them to &lt;a href=&#34;https://github.com/tesseract-ocr/tesseract&#34;&gt;Tesseract&lt;/a&gt; to see if I could get better results. After a bit of tweaking, things were looking pretty good. But when I came to compile all the data, I realised there was a potential problem matching the PDF pages to the images available through IIIF. I found one case where some pages were missing from the PDF, and another couple where the page order was different.&lt;/p&gt;
&lt;p&gt;As I was looking around for a solution, I realised that those JSON files I downloaded to get the PDF identifiers also included links to &lt;a href=&#34;https://en.wikipedia.org/wiki/Analyzed_Layout_and_Text_Object&#34;&gt;ALTO XML&lt;/a&gt; files that contain all  the original OCR data (before it got mangled by the PDF formatting). There was one ALTO file for every page. Even better, the JSON linked the identifiers for the text and the image together – no more page mismatches!&lt;/p&gt;
&lt;h3 id=&#34;downloading-the-alto-files&#34;&gt;Downloading the ALTO files&lt;/h3&gt;
&lt;p&gt;Let&amp;rsquo;s start this again shall we. After wasting several days futzing about with the PDFs, I decided to download all the ALTO files and extract the text from them. As I downloaded each XML file, I also grabbed the corresponding image identifier from the JSON and included both identifiers in the file name for safe keeping.&lt;/p&gt;
&lt;p&gt;The ALTO files break the text down by block, line, and word. To extract the text, I just looped through every line, joining the words back together as a string, and writing the result to a new text file – one for each page.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s worth noting that the ALTO files include &lt;em&gt;all&lt;/em&gt; the positional data generated by the OCR process, so you have the size and position of every word on every page. I just pulled out the text, but there are many more interesting things you could do&amp;hellip;&lt;/p&gt;
&lt;h3 id=&#34;assembling-and-publishing-the-database&#34;&gt;Assembling and publishing the database&lt;/h3&gt;
&lt;p&gt;From here on everything pretty much followed the pattern of the NSW and Tasmanian directories. I looped through each volume, page, and line of text, adding the text and metadata to a SQLite database using &lt;a href=&#34;https://sqlite-utils.datasette.io/en/stable/&#34;&gt;sqlite_utils&lt;/a&gt;. I then indexed the text for full-text searching. At the same time I populated a metadata file with titles, urls, and few configuration details. The metadata file is used by &lt;a href=&#34;https://datasette.io/&#34;&gt;Datasette&lt;/a&gt; to fill in parts of the interface.&lt;/p&gt;
&lt;p&gt;I made some minor changes to the Datasette template I used for the other directories. In particular, I had to update the urls that loaded the &lt;a href=&#34;https://iiif.io&#34;&gt;IIIF&lt;/a&gt; images into the &lt;a href=&#34;https://openseadragon.github.io&#34;&gt;OpenSeadragon viewer&lt;/a&gt;. But it mostly just worked. It&amp;rsquo;s so nice to be able to reuse existing patterns!&lt;/p&gt;
&lt;p&gt;Finally, I used &lt;a href=&#34;https://docs.datasette.io/en/stable/publish.html&#34;&gt;Datasette&amp;rsquo;s &lt;code&gt;publish&lt;/code&gt; command&lt;/a&gt; to push everything to Google Cloudrun. The final database contains details of more than 50,000 pages, and over 19 million lines of text! It weighs in at about 1.7gb. The Cloudrun service will &amp;lsquo;scale to zero&amp;rsquo; when not in use. This saves some money and resources, but means it can take a little while to spin up. Once it&amp;rsquo;s loaded, it&amp;rsquo;s very fast. My &lt;a href=&#34;https://updates.timsherratt.org/2022/09/15/from-pdfs-to.html&#34;&gt;original post on the Tasmanian directories&lt;/a&gt; included a little note on costs, if you&amp;rsquo;re interested.&lt;/p&gt;
&lt;h2 id=&#34;more-information&#34;&gt;More information&lt;/h2&gt;
&lt;p&gt;The notebooks I used are on GitHub:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_sands_and_mac_pdfs.ipynb&#34;&gt;Download Sands and Mac PDFs and OCR text&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/load_sands_and_mac_into_datasette.ipynb&#34;&gt;Load data from the Sands and Mac directories into an SQLite database (for use with Datasette)&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Here are some posts about the NSW and Tasmanian directories:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2022/09/01/making-nsw-postal.html&#34;&gt;Making NSW Postal Directories (and other digitised directories) easier to search with the GLAM Workbench and Datasette&lt;/a&gt; (September 2022)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2022/09/15/from-pdfs-to.html&#34;&gt;From 48 PDFs to one searchable database – opening up the Tasmanian Post Office Directories with the GLAM Workbench&lt;/a&gt; (September 2022)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2024/09/26/wheres-missing-volume.html&#34;&gt;Where&amp;rsquo;s 1920? Missing volume added to Tasmanian Post Office Directories!&lt;/a&gt; (September 2024)&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://updates.timsherratt.org/2024/11/21/six-more-volumes.html&#34;&gt;Six more volumes added to the searchable database of Tasmanian Post Office Directories!&lt;/a&gt; (November 2024)&lt;/li&gt;
&lt;/ul&gt;
</description>
      <source:markdown>In the fortnight I spent onsite at the State Library of Victoria, &#39;Sands &amp; Mac&#39; was mentioned many times. And no wonder. The [Sands &amp; McDougall&#39;s directories](https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;collectionId=81213035910007636) are a goldmine for anyone researching family, local, or social history. They list thousands of names and addresses, enabling you to find individuals, and explore changing land use over time. When people ask the SLV&#39;s librarians, &#39;What can you tell me about the history of my house?&#39;, Sands &amp; Mac is one of the first resources consulted.

The SLV has digitised [24 volumes of Sands &amp; Mac](https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;collectionId=81213035910007636), one every five years from 1860 to 1974. You can browse the contents of each volume in the SLV image viewer, using the partial contents listing to help you find your way to sections of interest. To search the full text content you need to use the PDF version, either in the built-in viewer, or by downloading the PDF. There&#39;s a [handy guide to using Sands &amp; Mac](https://blogs.slv.vic.gov.au/tips-and-tricks/collection-discovery-tips-sands-mcdougalls-directories/) that explains the options.

**However, there&#39;s currently no way of searching across all 24 volumes, so as part of my residency at the SLV LAB, I thought I&#39;d make one!**

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac.png&#34; width=&#34;600&#34; height=&#34;310&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/&#34;&gt;&lt;b&gt;Try it now!&lt;/b&gt;&lt;/a&gt;&lt;/figcaption&gt;&lt;/figure&gt;

My [new Sands &amp; Mac database](https://glam-workbench.net/state-library-victoria/sands-macdougall-directories/) follows the pattern I&#39;ve used previously to create fully-searchable versions of the [NSW Post Office directories](https://glam-workbench.net/trove-journals/nsw-post-office-directories/), [Sydney telephone directories](https://glam-workbench.net/trove-journals/sydney-telephone-directories/), and [Tasmanian Post Office directories](https://glam-workbench.net/tasmanian-post-office-directories/). Every line of text is saved to a database, so a single query searches for entries across all volumes. You can also use advanced search features like wildcards and boolean operators.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac-search.png&#34; width=&#34;600&#34; height=&#34;543&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Search across all 24 volumes!&lt;/figcaption&gt;&lt;/figure&gt;

Once you&#39;ve found a relevant entry you can view it in context, alongside a zoomable image of the page. You can even use Zotero to save individual entries to your own research database. [This blog post](https://chineseaustralia.org/from-the-archive-uncovering-the-everyday-heritage-of-chinese-tasmanians/) from the Everyday Heritage project describes how the Tasmanian directories have been used to map Tasmania&#39;s Chinese population.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/sands-and-mac-entry.png&#34; width=&#34;600&#34; height=&#34;370&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;View each entry in context! (Here&#39;s my Dad building his first house in Beaumaris in the 1950s.)&lt;/figcaption&gt;&lt;/figure&gt;

There&#39;s still a few things I&#39;d like to try, such as making use of the table of contents information for each volume. I&#39;d also like to create some additional entry points to take users directly to listings for individual suburbs (maybe even streets!). Each volume has a directory of suburbs, so it would be a matter of extracting and cleaning the data and linking the entries to digitised pages. Certainly possible, but I don&#39;t think I&#39;ll have time to get it all done before the end of my residency. Perhaps I&#39;ll try to get at least one volume done to demonstrate how it might work, and the value it would add. As I was writing this blog post I also realised there&#39;s [a dataset of businesses](https://www.environment.vic.gov.au/sustainability/victoria-unearthed/about-the-data/sands-and-mcdougall) extracted from the Sands &amp; Mac, so I need to think about how I can use that as well!

## Technical information follows...

I&#39;ve documented the process I used to create fully-searchable versions of the [Tasmanian](https://glam-workbench.net/libraries-tasmania/) and [NSW directories](https://glam-workbench.net/trove-journals/create-text-db-indexed-by-line/)  in the GLAM Workbench. I followed a similar method for Sands and Mac, though with a few dead-ends and discoveries along the way.

### Downloading the PDFs

I assumed that it would be easiest to work from the PDF versions of each volume, as I&#39;d done for Tasmania. So I set about finding a way to download them all. There&#39;s only 24 volumes, so I *could* have downloaded them manually, but where&#39;s the fun in that?

I started with a CSV file listing the Sands &amp; Mac volumes that I downloaded from the catalogue. This gave me the Alma identifiers for each volume. To download the PDFs I needed two more identifiers, the `IE` identifier assigned to each digitised item, and a file identifier that points to the PDF version of the item. The `IE` identifier can be extracted from the item&#39;s MARC record, as I described in [my post on exploring urls](https://updates.timsherratt.org/2025/09/23/exploring-slv-urls.html). The PDF file identifier was a bit more difficult to track down. The PDF links in the image viewer are generated dynamically, so the data had to be coming from somewhere. Eventually I found that the viewer loaded a JSON file with all sorts of useful metadata in it! 

The url to download the JSON file is: `https://viewerapi.slv.vic.gov.au/?entity=[IE identifier]&amp;dc_arrays=1`. In the `summary` section I found identifiers for `small_pdf` and `master_pdf`. 

I could then use these identifiers to construct urls to download the PDFs themselves: `https://rosetta.slv.vic.gov.au/delivery/DeliveryManagerServlet?dps_func=stream&amp;dps_pid=[PDF id]`

Once I had the PDFs I used [PyMuPDF](https://github.com/pymupdf/PyMuPDF) to extract all the text and images. As I suspected the text wasn&#39;t really fit for purpose. The OCR was ok, but the column structures were a mess. Because I wanted to index each entry individually, it was important to try and get the columns represented as accurately as possible. The images in the small PDFs were already bitonal, so I started feeding them to [Tesseract](https://github.com/tesseract-ocr/tesseract) to see if I could get better results. After a bit of tweaking, things were looking pretty good. But when I came to compile all the data, I realised there was a potential problem matching the PDF pages to the images available through IIIF. I found one case where some pages were missing from the PDF, and another couple where the page order was different.

As I was looking around for a solution, I realised that those JSON files I downloaded to get the PDF identifiers also included links to [ALTO XML](https://en.wikipedia.org/wiki/Analyzed_Layout_and_Text_Object) files that contain all  the original OCR data (before it got mangled by the PDF formatting). There was one ALTO file for every page. Even better, the JSON linked the identifiers for the text and the image together – no more page mismatches!

### Downloading the ALTO files

Let&#39;s start this again shall we. After wasting several days futzing about with the PDFs, I decided to download all the ALTO files and extract the text from them. As I downloaded each XML file, I also grabbed the corresponding image identifier from the JSON and included both identifiers in the file name for safe keeping.

The ALTO files break the text down by block, line, and word. To extract the text, I just looped through every line, joining the words back together as a string, and writing the result to a new text file – one for each page.

It&#39;s worth noting that the ALTO files include *all* the positional data generated by the OCR process, so you have the size and position of every word on every page. I just pulled out the text, but there are many more interesting things you could do...

### Assembling and publishing the database

From here on everything pretty much followed the pattern of the NSW and Tasmanian directories. I looped through each volume, page, and line of text, adding the text and metadata to a SQLite database using [sqlite_utils](https://sqlite-utils.datasette.io/en/stable/). I then indexed the text for full-text searching. At the same time I populated a metadata file with titles, urls, and few configuration details. The metadata file is used by [Datasette](https://datasette.io/) to fill in parts of the interface.

I made some minor changes to the Datasette template I used for the other directories. In particular, I had to update the urls that loaded the [IIIF](https://iiif.io) images into the [OpenSeadragon viewer](https://openseadragon.github.io). But it mostly just worked. It&#39;s so nice to be able to reuse existing patterns!

Finally, I used [Datasette&#39;s `publish` command](https://docs.datasette.io/en/stable/publish.html) to push everything to Google Cloudrun. The final database contains details of more than 50,000 pages, and over 19 million lines of text! It weighs in at about 1.7gb. The Cloudrun service will &#39;scale to zero&#39; when not in use. This saves some money and resources, but means it can take a little while to spin up. Once it&#39;s loaded, it&#39;s very fast. My [original post on the Tasmanian directories](https://updates.timsherratt.org/2022/09/15/from-pdfs-to.html) included a little note on costs, if you&#39;re interested.

## More information

The notebooks I used are on GitHub:

- [Download Sands and Mac PDFs and OCR text](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/download_sands_and_mac_pdfs.ipynb)
- [Load data from the Sands and Mac directories into an SQLite database (for use with Datasette)](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/load_sands_and_mac_into_datasette.ipynb)

Here are some posts about the NSW and Tasmanian directories:

- [Making NSW Postal Directories (and other digitised directories) easier to search with the GLAM Workbench and Datasette](https://updates.timsherratt.org/2022/09/01/making-nsw-postal.html) (September 2022)
- [From 48 PDFs to one searchable database – opening up the Tasmanian Post Office Directories with the GLAM Workbench](https://updates.timsherratt.org/2022/09/15/from-pdfs-to.html) (September 2022)
- [Where&#39;s 1920? Missing volume added to Tasmanian Post Office Directories!](https://updates.timsherratt.org/2024/09/26/wheres-missing-volume.html) (September 2024)
- [Six more volumes added to the searchable database of Tasmanian Post Office Directories!](https://updates.timsherratt.org/2024/11/21/six-more-volumes.html) (November 2024)


</source:markdown>
    </item>
    
    <item>
      <title>Turning the SLV&#39;s maps into data with Allmaps and some GLAM plumbing</title>
      <link>https://updates.timsherratt.org/2025/11/04/turning-the-slvs-maps-into.html</link>
      <pubDate>Tue, 04 Nov 2025 14:02:53 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/11/04/turning-the-slvs-maps-into.html</guid>
      <description>&lt;p&gt;I often describe what I do as GLAM data plumbing. Most of the time I&amp;rsquo;m not creating new tools, I&amp;rsquo;m figuring out what data is available and how I can connect it up to &lt;em&gt;existing&lt;/em&gt; tools. It&amp;rsquo;s rarely straightforward, but if I can get all the pipes connected and data flowing in the right direction, suddenly new things become possible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Things like turning all the State Library of Victoria&amp;rsquo;s digitised maps into data.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve just &lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;created a workflow&lt;/a&gt; that uses &lt;a href=&#34;https://allmaps.org&#34;&gt;AllMaps&lt;/a&gt; and &lt;a href=&#34;https://iiif.io/&#34;&gt;IIIF&lt;/a&gt; to georeference the SLV&amp;rsquo;s digitised maps. There&amp;rsquo;s some technical details below, but the idea is pretty simple. A userscript links the SLV image viewer to Allmaps – so you just click on a button, and the digitised map opens, ready for georeferencing.&lt;/p&gt;
&lt;p&gt;Why is this useful? Georeferencing relates a digitised map to real world geography. It describes the map&amp;rsquo;s position and extent using geospatial coordinates – turning historic documents into geospatial data that can be indexed, visualised and manipulated. Georeferencing opens digitised maps to new research uses.&lt;/p&gt;
&lt;p&gt;So, how many maps we can georeference before my residency finishes in December? Hundreds? Thousands? If you like maps and want to help, head to &lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;the documentation page&lt;/a&gt; to find out how to get started. And if you want to see how things are progressing, have a look at &lt;a href=&#34;https://wragge.github.io/slv-allmaps/dashboard.html&#34;&gt;the project dashboard&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/slv-allmaps-docs.png&#34; width=&#34;600&#34; height=&#34;466&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;View the documentation&lt;/a&gt; to get started&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;A few technical details follow&amp;hellip;&lt;/p&gt;
&lt;p&gt;Early on in my time as Creative Technologist-in-Residence at the State Library of Victoria, I started playing around with Allmaps for georeferencing digitised maps. It&amp;rsquo;s a great tool (really a suite of tools and standards) because instead of constructing a whole new platform it integrates with existing IIIF services. The SLV provides digitised images through IIIF, so I thought it should be possible to use Allmaps to georeference the SLV&amp;rsquo;s map collection.&lt;/p&gt;
&lt;p&gt;But I struck a problem that took some time to unravel. The IIIF urls in the SLV manifests include port numbers and that confused Allmaps. The manifests also sometimes contained references to image formats that weren&amp;rsquo;t actually accessible, generating errors when they were loaded. Hopefully these problems will be fixed by the SLV, but in the meantime I&amp;rsquo;ve created a proxy service that edits the manifest on the fly. The proxied urls can be loaded into the Allmaps Editor without errors. Pipes fixed, data flowing!&lt;/p&gt;
&lt;details&gt;
  &lt;summary&gt;Using the manifest proxy&lt;/summary&gt;
  &lt;p&gt;To generate a link to a proxied manifest, first grab the item&#39;s &lt;code&gt;IE&lt;/code&gt; identifier from the url of the digitised item viewer. For example, the identifier in this url &lt;code&gt;https://viewer.slv.vic.gov.au/?entity=IE15485265&amp;mode=browse&lt;/code&gt; is &lt;code&gt;IE15485265&lt;/code&gt;. Once you have the identifier, add it to the end of the url &lt;code&gt;https://wraggelabs.com/slv_iiif/&lt;/code&gt;. For example, &lt;a href=&#34;https://wraggelabs.com/slv_iiif/IE15485265&#34;&gt;https://wraggelabs.com/slv_iiif/IE15485265&lt;/a&gt;. You can then supply this url to the Allmaps editor.&lt;/p&gt;
&lt;/details&gt;
&lt;p&gt;But having to fiddle around with proxies didn&amp;rsquo;t make a great user experience. I needed some way of integrating the two services, so that a user could just click a button in the SLV website and start editing in Allmaps. Userscripts to the rescue!&lt;/p&gt;
&lt;p&gt;I wrote recently about &lt;a href=&#34;https://updates.timsherratt.org/2025/07/17/glam-hacking-with-userscripts.html&#34;&gt;hacking GLAM collection interfaces using userscripts&lt;/a&gt;. Since I started my residency at the SLV, I&amp;rsquo;ve also created a userscript to &lt;a href=&#34;https://gist.github.com/wragge/a37a4db854deffad956abc7bf918f6b0&#34;&gt;display the IIIF manifest url in the SLV image viewer&lt;/a&gt;, and run a Code Club workshop where we played around with &lt;a href=&#34;https://slides.com/wragge/slv-code-club&#34;&gt;an assortment of SLV website hacks&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;As in a number of these examples, the &lt;a href=&#34;https://gist.github.com/wragge/5680daaec4b4b34ed5537e6ff79559a2&#34;&gt;georeferencing userscript&lt;/a&gt; adds new features to the SLV website, but there&amp;rsquo;s a fair bit more going on under the hood. It runs automatically every time you load the SLV image viewer, and then:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;it checks the metadata of the digitised item to see it it&amp;rsquo;s a map (or something that contains maps, like an atlas or street directory)&lt;/li&gt;
&lt;li&gt;if it looks like a map, it generates an Allmaps identifier using the item&amp;rsquo;s IIIF manifest url and checks with Allmaps to see whether the item has already been georeferenced&lt;/li&gt;
&lt;li&gt;it adds a &amp;lsquo;Georeferencing&amp;rsquo; section to the page, with a button to georeference the item (or edit the existing georeferencing)&lt;/li&gt;
&lt;li&gt;if the item has already been georeferenced, it adds a button to view the item in the Allmaps Viewer, and embeds a live preview&lt;/li&gt;
&lt;/ul&gt;
&lt;details&gt;
    &lt;summary&gt;Accessing metadata&lt;/summary&gt;
    &lt;p&gt;
        The userscript gets the item metadata from a JSON file that&#39;s loaded by the image viewer. The JSON file includes a lot of extra, useful information about the digitised item. To access the JSON file, you just construct a url like this: &lt;code&gt;https://viewerapi.slv.vic.gov.au/?entity=[IE identifier]&amp;dc_arrays=1&lt;/code&gt;. The IE identifier is in the url of the image viewer.
    &lt;/p&gt;
&lt;/details&gt;
&lt;details&gt;
    &lt;summary&gt;Allmaps identifiers&lt;/summary&gt;
    &lt;p&gt;
        Allmaps creates its identifiers by hash encoding the IIIF urls. The userscript borrows some code from the &lt;a href=&#34;https://github.com/allmaps/allmaps/tree/main/packages/id&#34;&gt;Allmaps id module&lt;/a&gt; to generate the ids, then sends a HEAD request to the Allmaps API to see whether an entry for the current manifest exists.
    &lt;/p&gt;
&lt;/details&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/alv-allmaps-not-georeferenced.png&#34; width=&#34;600&#34; height=&#34;313&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Example of an item that hasn&#39;t been georeferenced yet&lt;/figcaption&gt;&lt;/figure&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/slv-allmaps-georeferenced.png&#34; width=&#34;600&#34; height=&#34;462&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Example of an item that has been georeferenced, displaying an embedded version of the Allmaps viewer&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;I&amp;rsquo;ve also created a GitHub repository to save copies of the data. Every two hours &lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/harvest_allmaps_data.ipynb&#34;&gt;this notebook&lt;/a&gt; is run to query the Allmaps API for newly georeferenced maps. These are added to a dataset which is saved in three formats:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.csv&#34;&gt;a CSV file&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps_datasette.csv&#34;&gt;a CSV file&lt;/a&gt; that includes thumbnails and links for &lt;a href=&#34;https://glam-workbench.net/datasette-lite/?csv=https%3A%2F%2Fgithub.com%2Fwragge%2Fslv-allmaps%2Fblob%2Fmain%2Fgeoreferenced_maps_datasette.csv&amp;amp;install=datasette-homepage-table&amp;amp;install=datasette-json-html&amp;amp;fts=manifest_title%2Cmap_title&#34;&gt;viewing in Datasette-Lite&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.geojson&#34;&gt;a GeoJSON file&lt;/a&gt;, that can be &lt;a href=&#34;https://geojson.io/#id=github:wragge/slv-allmaps/blob/main/georeferenced_maps.geojson&#34;&gt;viewed in services like geojson.io&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;At the same time, the data for each individual map is downloaded and saved as &lt;a href=&#34;https://github.com/wragge/slv-allmaps/tree/main/maps&#34;&gt;IIIF annotations&lt;/a&gt; (in JSON) and &lt;a href=&#34;https://github.com/wragge/slv-allmaps/tree/main/geojson&#34;&gt;GeoJSON&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Finally, &lt;a href=&#34;https://github.com/wragge/slv-allmaps/blob/main/allmaps_dashboard.ipynb&#34;&gt;this notebook&lt;/a&gt; is run to generate &lt;a href=&#34;https://wragge.github.io/slv-allmaps/dashboard.html&#34;&gt;a dashboard&lt;/a&gt; that provides an overview of the project&amp;rsquo;s progress.&lt;/p&gt;
&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/geo-dashboard.png&#34; width=&#34;600&#34; height=&#34;616&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;The project dashboard is updated every two hours&lt;/figcaption&gt;&lt;/figure&gt;
&lt;p&gt;One of the Allmaps developers described all my plumbing and workarounds as a &amp;lsquo;very cool lofi example of how you can set this up with little means&amp;rsquo;, and I think that&amp;rsquo;s pretty apt. It&amp;rsquo;s really just an experiment to demonstrate the possibilities, but by connecting up existing services it&amp;rsquo;s generating real data of long term value.&lt;/p&gt;
</description>
      <source:markdown>I often describe what I do as GLAM data plumbing. Most of the time I&#39;m not creating new tools, I&#39;m figuring out what data is available and how I can connect it up to *existing* tools. It&#39;s rarely straightforward, but if I can get all the pipes connected and data flowing in the right direction, suddenly new things become possible.

**Things like turning all the State Library of Victoria&#39;s digitised maps into data.**

I&#39;ve just [created a workflow](https://wragge.github.io/slv-allmaps/) that uses [AllMaps](https://allmaps.org) and [IIIF](https://iiif.io/) to georeference the SLV&#39;s digitised maps. There&#39;s some technical details below, but the idea is pretty simple. A userscript links the SLV image viewer to Allmaps – so you just click on a button, and the digitised map opens, ready for georeferencing. 

Why is this useful? Georeferencing relates a digitised map to real world geography. It describes the map&#39;s position and extent using geospatial coordinates – turning historic documents into geospatial data that can be indexed, visualised and manipulated. Georeferencing opens digitised maps to new research uses.

So, how many maps we can georeference before my residency finishes in December? Hundreds? Thousands? If you like maps and want to help, head to [the documentation page](https://wragge.github.io/slv-allmaps/) to find out how to get started. And if you want to see how things are progressing, have a look at [the project dashboard](https://wragge.github.io/slv-allmaps/dashboard.html).

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/slv-allmaps-docs.png&#34; width=&#34;600&#34; height=&#34;466&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;&lt;a href=&#34;https://wragge.github.io/slv-allmaps/&#34;&gt;View the documentation&lt;/a&gt; to get started&lt;/figcaption&gt;&lt;/figure&gt;

A few technical details follow...

Early on in my time as Creative Technologist-in-Residence at the State Library of Victoria, I started playing around with Allmaps for georeferencing digitised maps. It&#39;s a great tool (really a suite of tools and standards) because instead of constructing a whole new platform it integrates with existing IIIF services. The SLV provides digitised images through IIIF, so I thought it should be possible to use Allmaps to georeference the SLV&#39;s map collection.

But I struck a problem that took some time to unravel. The IIIF urls in the SLV manifests include port numbers and that confused Allmaps. The manifests also sometimes contained references to image formats that weren&#39;t actually accessible, generating errors when they were loaded. Hopefully these problems will be fixed by the SLV, but in the meantime I&#39;ve created a proxy service that edits the manifest on the fly. The proxied urls can be loaded into the Allmaps Editor without errors. Pipes fixed, data flowing!

&lt;details&gt;
  &lt;summary&gt;Using the manifest proxy&lt;/summary&gt;
  &lt;p&gt;To generate a link to a proxied manifest, first grab the item&#39;s &lt;code&gt;IE&lt;/code&gt; identifier from the url of the digitised item viewer. For example, the identifier in this url &lt;code&gt;https://viewer.slv.vic.gov.au/?entity=IE15485265&amp;mode=browse&lt;/code&gt; is &lt;code&gt;IE15485265&lt;/code&gt;. Once you have the identifier, add it to the end of the url &lt;code&gt;https://wraggelabs.com/slv_iiif/&lt;/code&gt;. For example, &lt;a href=&#34;https://wraggelabs.com/slv_iiif/IE15485265&#34;&gt;https://wraggelabs.com/slv_iiif/IE15485265&lt;/a&gt;. You can then supply this url to the Allmaps editor.&lt;/p&gt;
&lt;/details&gt;

But having to fiddle around with proxies didn&#39;t make a great user experience. I needed some way of integrating the two services, so that a user could just click a button in the SLV website and start editing in Allmaps. Userscripts to the rescue!

I wrote recently about [hacking GLAM collection interfaces using userscripts](https://updates.timsherratt.org/2025/07/17/glam-hacking-with-userscripts.html). Since I started my residency at the SLV, I&#39;ve also created a userscript to [display the IIIF manifest url in the SLV image viewer](https://gist.github.com/wragge/a37a4db854deffad956abc7bf918f6b0), and run a Code Club workshop where we played around with [an assortment of SLV website hacks](https://slides.com/wragge/slv-code-club). 

As in a number of these examples, the [georeferencing userscript](https://gist.github.com/wragge/5680daaec4b4b34ed5537e6ff79559a2) adds new features to the SLV website, but there&#39;s a fair bit more going on under the hood. It runs automatically every time you load the SLV image viewer, and then:

- it checks the metadata of the digitised item to see it it&#39;s a map (or something that contains maps, like an atlas or street directory)
- if it looks like a map, it generates an Allmaps identifier using the item&#39;s IIIF manifest url and checks with Allmaps to see whether the item has already been georeferenced
- it adds a &#39;Georeferencing&#39; section to the page, with a button to georeference the item (or edit the existing georeferencing)
- if the item has already been georeferenced, it adds a button to view the item in the Allmaps Viewer, and embeds a live preview

&lt;details&gt;
    &lt;summary&gt;Accessing metadata&lt;/summary&gt;
    &lt;p&gt;
        The userscript gets the item metadata from a JSON file that&#39;s loaded by the image viewer. The JSON file includes a lot of extra, useful information about the digitised item. To access the JSON file, you just construct a url like this: &lt;code&gt;https://viewerapi.slv.vic.gov.au/?entity=[IE identifier]&amp;dc_arrays=1&lt;/code&gt;. The IE identifier is in the url of the image viewer.
    &lt;/p&gt;
&lt;/details&gt;

&lt;details&gt;
    &lt;summary&gt;Allmaps identifiers&lt;/summary&gt;
    &lt;p&gt;
        Allmaps creates its identifiers by hash encoding the IIIF urls. The userscript borrows some code from the &lt;a href=&#34;https://github.com/allmaps/allmaps/tree/main/packages/id&#34;&gt;Allmaps id module&lt;/a&gt; to generate the ids, then sends a HEAD request to the Allmaps API to see whether an entry for the current manifest exists.
    &lt;/p&gt;
&lt;/details&gt;

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/alv-allmaps-not-georeferenced.png&#34; width=&#34;600&#34; height=&#34;313&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Example of an item that hasn&#39;t been georeferenced yet&lt;/figcaption&gt;&lt;/figure&gt;

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/slv-allmaps-georeferenced.png&#34; width=&#34;600&#34; height=&#34;462&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;Example of an item that has been georeferenced, displaying an embedded version of the Allmaps viewer&lt;/figcaption&gt;&lt;/figure&gt;

I&#39;ve also created a GitHub repository to save copies of the data. Every two hours [this notebook](https://github.com/wragge/slv-allmaps/blob/main/harvest_allmaps_data.ipynb) is run to query the Allmaps API for newly georeferenced maps. These are added to a dataset which is saved in three formats:

- [a CSV file](https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.csv)
- [a CSV file](https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps_datasette.csv) that includes thumbnails and links for [viewing in Datasette-Lite](https://glam-workbench.net/datasette-lite/?csv=https%3A%2F%2Fgithub.com%2Fwragge%2Fslv-allmaps%2Fblob%2Fmain%2Fgeoreferenced_maps_datasette.csv&amp;install=datasette-homepage-table&amp;install=datasette-json-html&amp;fts=manifest_title%2Cmap_title)
- [a GeoJSON file](https://github.com/wragge/slv-allmaps/blob/main/georeferenced_maps.geojson), that can be [viewed in services like geojson.io](https://geojson.io/#id=github:wragge/slv-allmaps/blob/main/georeferenced_maps.geojson)

At the same time, the data for each individual map is downloaded and saved as [IIIF annotations](https://github.com/wragge/slv-allmaps/tree/main/maps) (in JSON) and [GeoJSON](https://github.com/wragge/slv-allmaps/tree/main/geojson).

Finally, [this notebook](https://github.com/wragge/slv-allmaps/blob/main/allmaps_dashboard.ipynb) is run to generate [a dashboard](https://wragge.github.io/slv-allmaps/dashboard.html) that provides an overview of the project&#39;s progress.

&lt;figure&gt;&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/geo-dashboard.png&#34; width=&#34;600&#34; height=&#34;616&#34; alt=&#34;&#34;&gt;&lt;figcaption&gt;The project dashboard is updated every two hours&lt;/figcaption&gt;&lt;/figure&gt;

One of the Allmaps developers described all my plumbing and workarounds as a &#39;very cool lofi example of how you can set this up with little means&#39;, and I think that&#39;s pretty apt. It&#39;s really just an experiment to demonstrate the possibilities, but by connecting up existing services it&#39;s generating real data of long term value.
</source:markdown>
    </item>
    
    <item>
      <title>Me at 63...</title>
      <link>https://updates.timsherratt.org/2025/11/03/me-at.html</link>
      <pubDate>Mon, 03 Nov 2025 16:55:17 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/11/03/me-at.html</guid>
      <description>&lt;p&gt;&lt;em&gt;Originally published in Pharos, newsletter of the Professional Historian&amp;rsquo;s Association (Vic &amp;amp; Tas), October-November 2025, in the &amp;lsquo;Member Profile&amp;rsquo; section.&lt;/em&gt;&lt;/p&gt;
&lt;h2 id=&#34;what-was-your-first-history-related-job-what-path-have-you-taken-since-then&#34;&gt;What was your first history related job? What path have you taken since then?&lt;/h2&gt;
&lt;p&gt;In the early 1990s I started working for a small self-funded organisation called the Australian Science Archives Project. Our mission was to preserve and raise awareness of Australia&amp;rsquo;s scientific past. When the web came along, we realised it provided an enormous opportunity to communicate history to the public. So I taught myself web development and created the first archives website in Australia. Since then my work has continued to explore what happens when we release GLAM collections into online spaces where people can see and use them differently.&lt;/p&gt;
&lt;h2 id=&#34;what-kind-of-work-have-you-done-what-are-you-working-on-now&#34;&gt;What kind of work have you done? What are you working on now?&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;ve had a range of jobs in the GLAM and university sectors. While &amp;lsquo;history&amp;rsquo; wasn&amp;rsquo;t often in my job title, I&amp;rsquo;ve always regarded myself as a historian first – whether I was coding, editing, writing, teaching, or managing, history was always the frame through which I understood my work. At the same time, I&amp;rsquo;ve maintained my own independent practice as a &amp;lsquo;historian and hacker&amp;rsquo;, developing tools and resources for other researchers, such as the &lt;a href=&#34;https://glam-workbench.net&#34;&gt;GLAM Workbench&lt;/a&gt;. Much of this work is unfunded, but by sharing it openly I&amp;rsquo;ve created new opportunities for collaboration. For example, I&amp;rsquo;m currently the &amp;lsquo;Creative Technologist-in-Residence&amp;rsquo; at the State Library of Victoria, bringing my years of GLAM hacking to bear on the Library&amp;rsquo;s place based collections.&lt;/p&gt;
&lt;h2 id=&#34;research-or-writing-what-do-you-enjoy-more-and-why&#34;&gt;Research or writing? (What do you enjoy more and why?)&lt;/h2&gt;
&lt;p&gt;Researching, or writing, or coding, or teaching, or outreaching (what is the correct verb?) – all have their joys and travails. For me, research is less about finding things in archives and libraries, and more about &lt;em&gt;how&lt;/em&gt; we find things in archives and libraries. I poke about in online collections to try and understand how they work, what they reveal, and what they hide. This often leads to the development of new tools, the writing of  documentation and blog posts, and sometimes even real, published articles. It&amp;rsquo;s a process that has consumed my life, for better or worse. Coding often slips into obsession when I have a gnarly problem to crack. Writing is a slog, but there&amp;rsquo;s nothing like the pleasure of a finely-turned sentence. Teaching is exhausting, but also exhilarating when you see the light bulb of understanding flick on.&lt;/p&gt;
&lt;h2 id=&#34;what-are-the-best-and-hardest-things-about-the-kind-of-work-you-do&#34;&gt;What are the best and hardest things about the kind of work you do?&lt;/h2&gt;
&lt;p&gt;The best thing, the absolute hands-down best thing, is hearing from people who use, or have benefited from the tools and resources that I&amp;rsquo;ve created. I make things to help researchers see and use GLAM collections in new ways, so finding out what they&amp;rsquo;ve been doing with my stuff always provides a much-needed jolt of inspiration.&lt;/p&gt;
&lt;p&gt;However, the flip side is that getting information about my tools and resources out to the people who might benefit most is hard and often frustrating work. I churn away in the social media mines, but people and organisations seem much more reluctant to share new work these days. There was a time (yeah, the good old days) when GLAM organisations actively engaged with researchers online, sharing the cool things people were doing with their collections. But not now. We all learn through the generosity of others, and I think its important that we find ways to support and enlarge the realm of generosity.&lt;/p&gt;
</description>
      <source:markdown>*Originally published in Pharos, newsletter of the Professional Historian&#39;s Association (Vic &amp; Tas), October-November 2025, in the &#39;Member Profile&#39; section.*

## What was your first history related job? What path have you taken since then?

In the early 1990s I started working for a small self-funded organisation called the Australian Science Archives Project. Our mission was to preserve and raise awareness of Australia&#39;s scientific past. When the web came along, we realised it provided an enormous opportunity to communicate history to the public. So I taught myself web development and created the first archives website in Australia. Since then my work has continued to explore what happens when we release GLAM collections into online spaces where people can see and use them differently.

## What kind of work have you done? What are you working on now?

I&#39;ve had a range of jobs in the GLAM and university sectors. While &#39;history&#39; wasn&#39;t often in my job title, I&#39;ve always regarded myself as a historian first – whether I was coding, editing, writing, teaching, or managing, history was always the frame through which I understood my work. At the same time, I&#39;ve maintained my own independent practice as a &#39;historian and hacker&#39;, developing tools and resources for other researchers, such as the [GLAM Workbench](https://glam-workbench.net). Much of this work is unfunded, but by sharing it openly I&#39;ve created new opportunities for collaboration. For example, I&#39;m currently the &#39;Creative Technologist-in-Residence&#39; at the State Library of Victoria, bringing my years of GLAM hacking to bear on the Library&#39;s place based collections.

## Research or writing? (What do you enjoy more and why?)

Researching, or writing, or coding, or teaching, or outreaching (what is the correct verb?) – all have their joys and travails. For me, research is less about finding things in archives and libraries, and more about *how* we find things in archives and libraries. I poke about in online collections to try and understand how they work, what they reveal, and what they hide. This often leads to the development of new tools, the writing of  documentation and blog posts, and sometimes even real, published articles. It&#39;s a process that has consumed my life, for better or worse. Coding often slips into obsession when I have a gnarly problem to crack. Writing is a slog, but there&#39;s nothing like the pleasure of a finely-turned sentence. Teaching is exhausting, but also exhilarating when you see the light bulb of understanding flick on.

## What are the best and hardest things about the kind of work you do?

The best thing, the absolute hands-down best thing, is hearing from people who use, or have benefited from the tools and resources that I&#39;ve created. I make things to help researchers see and use GLAM collections in new ways, so finding out what they&#39;ve been doing with my stuff always provides a much-needed jolt of inspiration.

However, the flip side is that getting information about my tools and resources out to the people who might benefit most is hard and often frustrating work. I churn away in the social media mines, but people and organisations seem much more reluctant to share new work these days. There was a time (yeah, the good old days) when GLAM organisations actively engaged with researchers online, sharing the cool things people were doing with their collections. But not now. We all learn through the generosity of others, and I think its important that we find ways to support and enlarge the realm of generosity.


</source:markdown>
    </item>
    
    <item>
      <title>Creating bounding boxes for parish maps in the SLV collection</title>
      <link>https://updates.timsherratt.org/2025/10/06/creating-bounding-boxes-for-parish.html</link>
      <pubDate>Mon, 06 Oct 2025 14:17:51 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/10/06/creating-bounding-boxes-for-parish.html</guid>
      <description>&lt;p&gt;The State Library of Victoria holds a collection of &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/search?query=series,exact,Parish%20maps%20of%20Victoria&amp;amp;vid=61SLV_INST:SLV&amp;amp;offset=0&#34;&gt;8,804 parish maps&lt;/a&gt;. As part of my residency at the SLV LAB, I&amp;rsquo;ve been poking around in the metadata.&lt;/p&gt;
&lt;p&gt;SLV staff have geocoded many of the parish maps using the &lt;a href=&#34;https://placenames.fsdf.org.au&#34;&gt;Composite Gazetteer of Australia&lt;/a&gt;, which provides coordinates for Victorian parishes and boroughs. These coordinates give us a point which should be roughly at the centre of each map, enabling us to visualise their locations and distribution. But how much area do they cover? To answer that question we need a bounding box that includes the coordinates of each corner of the map. We could create bounding boxes by using something like &lt;a href=&#34;https://allmaps.org&#34;&gt;AllMaps&lt;/a&gt; or &lt;a href=&#34;https://www.mapwarper.net&#34;&gt;MapWarper&lt;/a&gt; to georeference each individual map, but that&amp;rsquo;s going to take a while! As a quick and dirty alternative, I wondered if it was possible to generate approximate bounding boxes from the available metadata. It seems we can!&lt;/p&gt;
&lt;h2 id=&#34;the-metadata&#34;&gt;The metadata&lt;/h2&gt;
&lt;p&gt;There are three pieces of metadata we need to construct bounding boxes:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;the latitude and longitude of the centre point&lt;/li&gt;
&lt;li&gt;the size of the physical map&lt;/li&gt;
&lt;li&gt;the scale of the map (ie how the size of the map relates to real world)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The coordinates and scale can be included in a couple of different places in the map&amp;rsquo;s MARC record. The &lt;a href=&#34;https://www.loc.gov/marc/bibliographic/bd034.html&#34;&gt;&lt;code&gt;034&lt;/code&gt;&lt;/a&gt; field is specifically for &amp;lsquo;Coded Cartographic Mathematical Data&amp;rsquo;. The relevant subfields are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;$a&lt;/code&gt;: category of scale&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$b&lt;/code&gt;: constant ratio linear horizontal scale (this is the most likely type of scale)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$d&lt;/code&gt;: westernmost longitude&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$e&lt;/code&gt;: easternmost longitude&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$f&lt;/code&gt;: northernmost latitude&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$g&lt;/code&gt;: southernmost latitude&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;If the coordinates describe a point rather than a bounding box, then &lt;code&gt;$d&lt;/code&gt; and &lt;code&gt;$e&lt;/code&gt; will be the same, and &lt;code&gt;$f&lt;/code&gt; and &lt;code&gt;$g&lt;/code&gt; will be the same.&lt;/p&gt;
&lt;p&gt;String representations of coordinates and scale can be found in the &lt;code&gt;255&lt;/code&gt; field. The relevant subfields are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;$a&lt;/code&gt;: statement of scale, eg &lt;code&gt;Scale [ca. 1:90,000].&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;$c&lt;/code&gt;: statement of coordinates, eg &lt;code&gt;(E 142°18&#39;/S 37°33&#39;)&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The size of the map is recorded in &lt;a href=&#34;https://www.loc.gov/marc/bibliographic/bd300.html&#34;&gt;&lt;code&gt;300&lt;/code&gt;&lt;/a&gt; (physical description) field under the &lt;code&gt;$c&lt;/code&gt; (dimensions) subfield. For example: &lt;code&gt;on sheet 40 x 51 cm &lt;/code&gt;.&lt;/p&gt;
&lt;h2 id=&#34;the-method&#34;&gt;The method&lt;/h2&gt;
&lt;p&gt;I started with an existing dataset downloaded from the catalogue by SLV staff. This dataset included the scale and coordinate information in the &lt;code&gt;034&lt;/code&gt; field, and the coordinate string in &lt;code&gt;255$c&lt;/code&gt;. At first I didn&amp;rsquo;t realise that the &lt;code&gt;034&lt;/code&gt; held geo data, so I separately downloaded the scale information from &lt;code&gt;255:$a&lt;/code&gt; in each item&amp;rsquo;s MARC record (d&amp;rsquo;oh). If the maps were digitised, I also wanted their image identifiers so I could access them through the SLV&amp;rsquo;s IIIF service. The image id from the &lt;code&gt;956$e&lt;/code&gt; field of the MARC record can be used to construct an IIIF manifest url, so I extracted them as well.&lt;/p&gt;
&lt;p&gt;Once I had all the catalogue data, I had to make sure everything was in a format I could work with. The coordinates in the MARC records are recorded as degrees/minutes/seconds, so I had to convert them to decimal values. The scale factor needed to be an integer, and I needed to extract the height and width as integers from the dimensions field.&lt;/p&gt;
&lt;p&gt;I used &lt;a href=&#34;https://pypi.org/project/lat-lon-parser/&#34;&gt;lat_lon_parser&lt;/a&gt; to convert the coordinates to decimal, but needed a bit of regex string manipulation to get the values into a format that could be parsed. Regex also came to the rescue in getting the map dimensions. All the details are &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_parish_maps.ipynb&#34;&gt;in this notebook&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id=&#34;creating-bounding-boxes&#34;&gt;Creating bounding boxes&lt;/h3&gt;
&lt;p&gt;After some searching I found &lt;a href=&#34;https://stackoverflow.com/a/76910048&#34;&gt;this StackOverflow comment&lt;/a&gt; that described how to create a bounding box from a point, distance, and bearing. The point I already had, but the distance and bearing had to be calculated. Trigonometry to the rescue!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-box-trig.png&#34; width=&#34;600&#34; height=&#34;334&#34; alt=&#34;&#34;&gt;
&lt;p&gt;The distance from the point at the centre of the box to one of its corners is the hypotenuse of a right-angled triangle whose sides are equal to half the width and half the height of the map, and thanks to Pythagorus we know:&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-10-06-15-19-42.png&#34; width=&#34;579&#34; height=&#34;93&#34; alt=&#34;&#34;&gt;
&lt;p&gt;Once I had the distance in cm, I converted to inches, then multiplied by the scale factor, and finally converted the inches to miles. (It now occurs to me that there&amp;rsquo;s no need to convert to imperial measurements, but it doesn&amp;rsquo;t make any difference either way.)&lt;/p&gt;
&lt;p&gt;The bearing that points to the corner of the box is the angle inside the same right-angled triangle, so can be calculated using:&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-10-06-15-19-29.png&#34; width=&#34;407&#34; height=&#34;73&#34; alt=&#34;&#34;&gt;
&lt;p&gt;With the point of origin, distance, and bearing I could use &lt;a href=&#34;https://github.com/geopy/geopy&#34;&gt;geopy&lt;/a&gt; to calculate the corners of the bounding box!&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4&#34;&gt;&lt;code class=&#34;language-python&#34; data-lang=&#34;python&#34;&gt;&lt;span style=&#34;color:#f92672&#34;&gt;from&lt;/span&gt; geopy.distance &lt;span style=&#34;color:#f92672&#34;&gt;import&lt;/span&gt; geodesic

destination &lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt; geodesic(miles&lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt;distance)&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;destination(origin, bearing)
coords &lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt; destination&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;longitude, destination&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;latitude
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_parish_maps.ipynb&#34;&gt;See this notebook&lt;/a&gt; for the full details.&lt;/p&gt;
&lt;h2 id=&#34;limitations&#34;&gt;Limitations&lt;/h2&gt;
&lt;p&gt;Of course, this method is very rough and has a number of major limitations, in particular:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;only about 38% of the maps have point coordinates&lt;/li&gt;
&lt;li&gt;the point values don&amp;rsquo;t necessarily locate the centre of the map&lt;/li&gt;
&lt;li&gt;not all the maps are oriented towards north&lt;/li&gt;
&lt;li&gt;sometimes a parish includes multiple maps&lt;/li&gt;
&lt;li&gt;the size of the margin around the map will affect the accuracy of the bounding box&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;But despite these problems the results seem pretty good. To test this I &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_browser.ipynb&#34;&gt;created a notebook&lt;/a&gt; to overlay the digitised maps on a modern basemap using the bounding boxes. Here&amp;rsquo;s an example.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-overlay.png&#34; width=&#34;600&#34; height=&#34;471&#34; alt=&#34;Screenshot of a parish map of French Island overlaid on a modern basemap. The parish map is slightly offset to the north, but you can see that the size matches the modern map fairly well&#34;&gt;
&lt;p&gt;You can see the map is slightly offset (presumably due to the second problem listed above). But the size seems about right. Certainly good enough to use the bounding boxes in some exploratory analyses!&lt;/p&gt;
&lt;h2 id=&#34;visualising-the-results&#34;&gt;Visualising the results&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;ve saved the processed data as a &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_final.csv&#34;&gt;new dataset&lt;/a&gt;, and started playing around with a couple of ways of visualising the results. These are experiments, not discovery interfaces. But you can use them for a bit of exploration if you don&amp;rsquo;t mind a few bugs. They&amp;rsquo;re all in Jupyter notebooks that can be run &lt;a href=&#34;https://mybinder.org/v2/gh/StateLibraryVictoria-SLVLAB/geo-maps-residency/HEAD&#34;&gt;using the Binder service&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_browser.ipynb&#34;&gt;parish maps browser&lt;/a&gt; includes a dropdown list of parish maps with point coordinates.  Select a map and:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;if there&amp;rsquo;s a bounding box and an image identifier, the image of the parish map will be overlaid on the modern base map using the bounding box coordinates&lt;/li&gt;
&lt;li&gt;if there&amp;rsquo;s a bounding box, but no image identifier, a rectangle will be drawn on the base map showing the dimensions of the bounding box&lt;/li&gt;
&lt;li&gt;if there are point coordinates, but no bounding box, a marker will be placed on the base map&lt;/li&gt;
&lt;/ul&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-browser.png&#34; width=&#34;600&#34; height=&#34;429&#34; alt=&#34;Screenshot of a parish map of Mallacoota overlaid on a modern basemap. The opacity of the digitised map has been reduced making it easier to see how the two maps align. A popup is visible on the map, listing the basic metadata and including a link to the SLV catalogue.&#34;&gt;
&lt;p&gt;If the image of the map is displayed you can use the slider to adjust the opacity. Clicking on either the image, rectangle, or marker will display metadata about the parish map and a link to the SLV catalogue.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s also a &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_visualise_bounds.ipynb&#34;&gt;visualisation of all the bounding boxes&lt;/a&gt; overlaid on a modern base map.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-bounds.png&#34; width=&#34;600&#34; height=&#34;474&#34; alt=&#34;Screenshot of a modern digital map of Victoria overlaid with 3,000+ transparent blue rectangles, showing the bounds of parish maps. A couple of the maps seem to be in Bass Strait.&#34;&gt;
&lt;p&gt;As you move your mouse over the bounding boxes the titles are displayed on the map, and if you click on a bounding box the metadata is displayed beneath the map, including a link to the SLV catalogue.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s obvious from the image above that some of the coordinates must be wrong! Visualisation is a great way of finding problems with your data. I now need to work through the results, documenting the problems, and thinking about how to make best use of the data. More to come!&lt;/p&gt;
</description>
      <source:markdown>The State Library of Victoria holds a collection of [8,804 parish maps](https://find.slv.vic.gov.au/discovery/search?query=series,exact,Parish%20maps%20of%20Victoria&amp;vid=61SLV_INST:SLV&amp;offset=0). As part of my residency at the SLV LAB, I&#39;ve been poking around in the metadata.

SLV staff have geocoded many of the parish maps using the [Composite Gazetteer of Australia](https://placenames.fsdf.org.au), which provides coordinates for Victorian parishes and boroughs. These coordinates give us a point which should be roughly at the centre of each map, enabling us to visualise their locations and distribution. But how much area do they cover? To answer that question we need a bounding box that includes the coordinates of each corner of the map. We could create bounding boxes by using something like [AllMaps](https://allmaps.org) or [MapWarper](https://www.mapwarper.net) to georeference each individual map, but that&#39;s going to take a while! As a quick and dirty alternative, I wondered if it was possible to generate approximate bounding boxes from the available metadata. It seems we can!

## The metadata

There are three pieces of metadata we need to construct bounding boxes:

- the latitude and longitude of the centre point
- the size of the physical map
- the scale of the map (ie how the size of the map relates to real world)

The coordinates and scale can be included in a couple of different places in the map&#39;s MARC record. The [`034`](https://www.loc.gov/marc/bibliographic/bd034.html) field is specifically for &#39;Coded Cartographic Mathematical Data&#39;. The relevant subfields are:

- `$a`: category of scale
- `$b`: constant ratio linear horizontal scale (this is the most likely type of scale)
- `$d`: westernmost longitude
- `$e`: easternmost longitude
- `$f`: northernmost latitude
- `$g`: southernmost latitude

If the coordinates describe a point rather than a bounding box, then `$d` and `$e` will be the same, and `$f` and `$g` will be the same.

String representations of coordinates and scale can be found in the `255` field. The relevant subfields are:

- `$a`: statement of scale, eg `Scale [ca. 1:90,000].`
- `$c`: statement of coordinates, eg `(E 142°18&#39;/S 37°33&#39;)`

The size of the map is recorded in [`300`](https://www.loc.gov/marc/bibliographic/bd300.html) (physical description) field under the `$c` (dimensions) subfield. For example: `on sheet 40 x 51 cm `.

## The method

I started with an existing dataset downloaded from the catalogue by SLV staff. This dataset included the scale and coordinate information in the `034` field, and the coordinate string in `255$c`. At first I didn&#39;t realise that the `034` held geo data, so I separately downloaded the scale information from `255:$a` in each item&#39;s MARC record (d&#39;oh). If the maps were digitised, I also wanted their image identifiers so I could access them through the SLV&#39;s IIIF service. The image id from the `956$e` field of the MARC record can be used to construct an IIIF manifest url, so I extracted them as well.

Once I had all the catalogue data, I had to make sure everything was in a format I could work with. The coordinates in the MARC records are recorded as degrees/minutes/seconds, so I had to convert them to decimal values. The scale factor needed to be an integer, and I needed to extract the height and width as integers from the dimensions field.

I used [lat_lon_parser](https://pypi.org/project/lat-lon-parser/) to convert the coordinates to decimal, but needed a bit of regex string manipulation to get the values into a format that could be parsed. Regex also came to the rescue in getting the map dimensions. All the details are [in this notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_parish_maps.ipynb).

### Creating bounding boxes

After some searching I found [this StackOverflow comment](https://stackoverflow.com/a/76910048) that described how to create a bounding box from a point, distance, and bearing. The point I already had, but the distance and bearing had to be calculated. Trigonometry to the rescue!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-box-trig.png&#34; width=&#34;600&#34; height=&#34;334&#34; alt=&#34;&#34;&gt;

The distance from the point at the centre of the box to one of its corners is the hypotenuse of a right-angled triangle whose sides are equal to half the width and half the height of the map, and thanks to Pythagorus we know:

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-10-06-15-19-42.png&#34; width=&#34;579&#34; height=&#34;93&#34; alt=&#34;&#34;&gt;

Once I had the distance in cm, I converted to inches, then multiplied by the scale factor, and finally converted the inches to miles. (It now occurs to me that there&#39;s no need to convert to imperial measurements, but it doesn&#39;t make any difference either way.)

The bearing that points to the corner of the box is the angle inside the same right-angled triangle, so can be calculated using:

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-10-06-15-19-29.png&#34; width=&#34;407&#34; height=&#34;73&#34; alt=&#34;&#34;&gt;

With the point of origin, distance, and bearing I could use [geopy](https://github.com/geopy/geopy) to calculate the corners of the bounding box!

```python
from geopy.distance import geodesic

destination = geodesic(miles=distance).destination(origin, bearing)
coords = destination.longitude, destination.latitude
```

[See this notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/process_parish_maps.ipynb) for the full details.

## Limitations

Of course, this method is very rough and has a number of major limitations, in particular:

- only about 38% of the maps have point coordinates
- the point values don&#39;t necessarily locate the centre of the map
- not all the maps are oriented towards north
- sometimes a parish includes multiple maps
- the size of the margin around the map will affect the accuracy of the bounding box

But despite these problems the results seem pretty good. To test this I [created a notebook](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_browser.ipynb) to overlay the digitised maps on a modern basemap using the bounding boxes. Here&#39;s an example.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-overlay.png&#34; width=&#34;600&#34; height=&#34;471&#34; alt=&#34;Screenshot of a parish map of French Island overlaid on a modern basemap. The parish map is slightly offset to the north, but you can see that the size matches the modern map fairly well&#34;&gt;

You can see the map is slightly offset (presumably due to the second problem listed above). But the size seems about right. Certainly good enough to use the bounding boxes in some exploratory analyses!

## Visualising the results

I&#39;ve saved the processed data as a [new dataset](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_final.csv), and started playing around with a couple of ways of visualising the results. These are experiments, not discovery interfaces. But you can use them for a bit of exploration if you don&#39;t mind a few bugs. They&#39;re all in Jupyter notebooks that can be run [using the Binder service](https://mybinder.org/v2/gh/StateLibraryVictoria-SLVLAB/geo-maps-residency/HEAD).

The [parish maps browser](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_browser.ipynb) includes a dropdown list of parish maps with point coordinates.  Select a map and:

- if there&#39;s a bounding box and an image identifier, the image of the parish map will be overlaid on the modern base map using the bounding box coordinates
- if there&#39;s a bounding box, but no image identifier, a rectangle will be drawn on the base map showing the dimensions of the bounding box
- if there are point coordinates, but no bounding box, a marker will be placed on the base map

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-browser.png&#34; width=&#34;600&#34; height=&#34;429&#34; alt=&#34;Screenshot of a parish map of Mallacoota overlaid on a modern basemap. The opacity of the digitised map has been reduced making it easier to see how the two maps align. A popup is visible on the map, listing the basic metadata and including a link to the SLV catalogue.&#34;&gt;

If the image of the map is displayed you can use the slider to adjust the opacity. Clicking on either the image, rectangle, or marker will display metadata about the parish map and a link to the SLV catalogue.

There&#39;s also a [visualisation of all the bounding boxes](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency/blob/main/parish_maps_visualise_bounds.ipynb) overlaid on a modern base map. 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/parish-maps-bounds.png&#34; width=&#34;600&#34; height=&#34;474&#34; alt=&#34;Screenshot of a modern digital map of Victoria overlaid with 3,000+ transparent blue rectangles, showing the bounds of parish maps. A couple of the maps seem to be in Bass Strait.&#34;&gt;

As you move your mouse over the bounding boxes the titles are displayed on the map, and if you click on a bounding box the metadata is displayed beneath the map, including a link to the SLV catalogue.

It&#39;s obvious from the image above that some of the coordinates must be wrong! Visualisation is a great way of finding problems with your data. I now need to work through the results, documenting the problems, and thinking about how to make best use of the data. More to come!
</source:markdown>
    </item>
    
    <item>
      <title>Exploring SLV urls</title>
      <link>https://updates.timsherratt.org/2025/09/23/exploring-slv-urls.html</link>
      <pubDate>Tue, 23 Sep 2025 16:22:45 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/09/23/exploring-slv-urls.html</guid>
      <description>&lt;p&gt;I like urls. They take you places. And if you know how to read them, they can tell you things about the systems that created them.
One of the first things I did when I started &lt;a href=&#34;https://updates.timsherratt.org/2025/09/22/creative-technologistinresidence-at-the-state.html&#34;&gt;my residency at SLV LAB&lt;/a&gt;, was to try and understand how their collection urls work. There&amp;rsquo;s a couple of well-worn methods I use when digging into a new site.&lt;/p&gt;
&lt;p&gt;The first is url hacking – this involves fiddling around with the parameters in a url and submitting the result to see what happens. The Trove Data Guide includes &lt;a href=&#34;https://tdg.glam-workbench.net/understanding-search/search-hacks.html&#34;&gt;some examples of hacking Trove urls&lt;/a&gt; to change the delivery of search results.&lt;/p&gt;
&lt;p&gt;The second method involves opening up the developer console in your web browser and watching the activity in the network tab as you click on links. This tells you where the information that gets loaded into your browser actually comes from – sometimes exposing handy urls that you can use to shortcut access to useful data.&lt;/p&gt;
&lt;h2 id=&#34;permalinks&#34;&gt;Permalinks&lt;/h2&gt;
&lt;p&gt;The SLV uses Primo for its public-facing catalogue, as well as other systems such as Rosetta and IIIF to deliver digitised content. I&amp;rsquo;d noticed that &lt;a href=&#34;https://www.zotero.org&#34;&gt;Zotero&lt;/a&gt; gets some useful data from the catalogue using the default &amp;lsquo;Primo 2018&amp;rsquo; translator, however, important things like the item url aren&amp;rsquo;t captured. The problem is that Primo&amp;rsquo;s &amp;lsquo;permalinks&amp;rsquo; are generated as required by a browser click – they&amp;rsquo;re not embedded anywhere on the page. This makes it hard to Zotero to grab them. So I started wondering how Zotero could construct short, persistent(ish) links to items.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s a link to an item in Primo: &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;context=L&amp;amp;docid=alma9941325055707636&#34;&gt;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;context=L&amp;amp;docid=alma9941325055707636&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;It looks pretty long and messy, but if you start deleting parameters and resubmitting, you&amp;rsquo;ll find that only two parameters are essential, &lt;code&gt;vid&lt;/code&gt; and &lt;code&gt;docid&lt;/code&gt;. This means we can rewrite the url as: &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;docid=alma9941325055707636&#34;&gt;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;docid=alma9941325055707636&lt;/a&gt; Much nicer.&lt;/p&gt;
&lt;p&gt;The &amp;lsquo;permalink&amp;rsquo; for the same item is: &lt;a href=&#34;https://find.slv.vic.gov.au/permalink/61SLV_INST/1sev8ar/alma9941325055707636&#34;&gt;https://find.slv.vic.gov.au/permalink/61SLV_INST/1sev8ar/alma9941325055707636&lt;/a&gt; If you look closely at the url path and compare it to the example above you&amp;rsquo;ll see the path is constructed from &lt;code&gt;/vid/[some other id]/docid&lt;/code&gt;. One of the librarians explained to me that the other identifier in the permalink is an encoding of the view type, but given that the &amp;lsquo;fulldisplay&amp;rsquo; view is the default, we don&amp;rsquo;t really need it. So the shortened url seems fine for use in Zotero and is easy to generate from the current url. Nice.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s also worth noting that the &lt;code&gt;vid&lt;/code&gt; value doesn&amp;rsquo;t seem to change, so to construct catalogue urls in your code, all you really need is the ALMA identifier that&amp;rsquo;s in the &lt;code&gt;docid&lt;/code&gt; parameter.&lt;/p&gt;
&lt;h2 id=&#34;structured-data&#34;&gt;Structured data&lt;/h2&gt;
&lt;p&gt;Item pages in Primo include a link labelled &amp;lsquo;Display source record&amp;rsquo;. If you click on this you&amp;rsquo;re taken to a representation of the item&amp;rsquo;s metadata in MARC. Here&amp;rsquo;s what the urls look like: &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/sourceRecord?vid=61SLV_INST%3ASLV&amp;amp;docId=alma9941325055707636&amp;amp;recordOwner=61SLV_INST&#34;&gt;https://find.slv.vic.gov.au/discovery/sourceRecord?vid=61SLV_INST%3ASLV&amp;amp;docId=alma9941325055707636&amp;amp;recordOwner=61SLV_INST&lt;/a&gt; Notice that the &amp;lsquo;fulldisplay&amp;rsquo; in the url path above has changed to &amp;lsquo;sourceRecord&amp;rsquo;. There&amp;rsquo;s also a new &lt;code&gt;recordOwner&lt;/code&gt; parameter, but it seems you can delete this and still get the same result.&lt;/p&gt;
&lt;p&gt;Having access to the MARC record is handy, because it delivers the metadata in a simple, structured plain text format. But while the &amp;lsquo;source record&amp;rsquo; page looks like a plain text file, it&amp;rsquo;s actually a HTML page that embeds a plain text record. If you open up the network tab of your browser&amp;rsquo;s developer console and reload the &amp;lsquo;source record&amp;rsquo; page, you&amp;rsquo;ll see a different url is loaded under the hood: &lt;a href=&#34;https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;amp;vid=61SLV_INST:SLV&amp;amp;recordOwner=61SLV_INST&amp;amp;lang=en&#34;&gt;https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;amp;vid=61SLV_INST:SLV&amp;amp;recordOwner=61SLV_INST&amp;amp;lang=en&lt;/a&gt; See how the url path has changed from &lt;code&gt;/discovery/&lt;/code&gt; to &lt;code&gt;/primaws/rest/pub&lt;/code&gt;? This url &lt;em&gt;does&lt;/em&gt; deliver a plain text version of the MARC record.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-23-15-21-05.png&#34; width=&#34;600&#34; height=&#34;215&#34; alt=&#34;&#34;&gt;
&lt;p&gt;Once you have the plain text version you can parse the contents to extract the structured data. There are tools that can probably do this automatically, but it&amp;rsquo;s also pretty easy using regular expressions. Here&amp;rsquo;s an example of some code I used to parse map records.&lt;/p&gt;
&lt;div class=&#34;highlight&#34;&gt;&lt;pre tabindex=&#34;0&#34; style=&#34;color:#f8f8f2;background-color:#272822;-moz-tab-size:4;-o-tab-size:4;tab-size:4&#34;&gt;&lt;code class=&#34;language-python&#34; data-lang=&#34;python&#34;&gt;&lt;span style=&#34;color:#66d9ef&#34;&gt;def&lt;/span&gt; &lt;span style=&#34;color:#a6e22e&#34;&gt;get_marc_value&lt;/span&gt;(marc, tag, subfield):
    &lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;&amp;#34;&amp;#34;
&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    Gets the value of a tag/subfield from a text version of an item&amp;#39;s MARC record.
&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;    &amp;#34;&amp;#34;&amp;#34;&lt;/span&gt;
    &lt;span style=&#34;color:#66d9ef&#34;&gt;try&lt;/span&gt;:
        tag &lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt; re&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;search(&lt;span style=&#34;color:#e6db74&#34;&gt;rf&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;^&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;{&lt;/span&gt;tag&lt;span style=&#34;color:#e6db74&#34;&gt;}&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;\t.+&amp;#34;&lt;/span&gt;, marc, re&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;M)&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;group(&lt;span style=&#34;color:#ae81ff&#34;&gt;0&lt;/span&gt;)
        subfield &lt;span style=&#34;color:#f92672&#34;&gt;=&lt;/span&gt; re&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;search(&lt;span style=&#34;color:#e6db74&#34;&gt;rf&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34;\$&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;{&lt;/span&gt;subfield&lt;span style=&#34;color:#e6db74&#34;&gt;}&lt;/span&gt;&lt;span style=&#34;color:#e6db74&#34;&gt;([^\$]+)&amp;#34;&lt;/span&gt;, tag)&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;group(&lt;span style=&#34;color:#ae81ff&#34;&gt;1&lt;/span&gt;)
    &lt;span style=&#34;color:#66d9ef&#34;&gt;except&lt;/span&gt; &lt;span style=&#34;color:#a6e22e&#34;&gt;AttributeError&lt;/span&gt;:
        &lt;span style=&#34;color:#66d9ef&#34;&gt;return&lt;/span&gt; &lt;span style=&#34;color:#66d9ef&#34;&gt;None&lt;/span&gt;
    &lt;span style=&#34;color:#66d9ef&#34;&gt;return&lt;/span&gt; subfield&lt;span style=&#34;color:#f92672&#34;&gt;.&lt;/span&gt;strip(&lt;span style=&#34;color:#e6db74&#34;&gt;&amp;#34; .,&amp;#34;&lt;/span&gt;)
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;You can also access a JSON representation of the record by adding the parameter &lt;code&gt;&amp;amp;showPnx=true&lt;/code&gt; to the catalogue url: &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;context=L&amp;amp;docid=alma9941325055707636&amp;amp;showPnx=true&#34;&gt;https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;context=L&amp;amp;docid=alma9941325055707636&amp;amp;showPnx=true&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Once again, this is a JSON representation embedded in a web page. Using the same developer console trick, you can identify the direct url is: &lt;a href=&#34;https://find.slv.vic.gov.au/primaws/rest/pub/pnxs/L/alma9941325055707636?vid=61SLV_INST:SLV&amp;amp;lang=en&amp;amp;search_scope=slv_local&amp;amp;showPnx=true&amp;amp;lang=en&#34;&gt;https://find.slv.vic.gov.au/primaws/rest/pub/pnxs/L/alma9941325055707636?vid=61SLV_INST:SLV&amp;amp;lang=en&amp;amp;search_scope=slv_local&amp;amp;showPnx=true&amp;amp;lang=en&lt;/a&gt; You should be able to parse the response from this url as JSON and use it in your code. I think the Zotero translator makes use of this &lt;code&gt;pnx&lt;/code&gt; data.&lt;/p&gt;
&lt;p&gt;If you want to download the MARC or JSON representations in your code, all you really need is the &lt;code&gt;alma&lt;/code&gt; identifier. Just use it to construct one of the direct urls, such as this: &lt;a href=&#34;https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;amp;vid=61SLV_INST:SLV&#34;&gt;https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;amp;vid=61SLV_INST:SLV&lt;/a&gt; The &lt;code&gt;recordOwner&lt;/code&gt; and &lt;code&gt;lang&lt;/code&gt; parameters are not needed, and the &lt;code&gt;vid&lt;/code&gt; parameter doesn&amp;rsquo;t change.&lt;/p&gt;
&lt;p&gt;Librarians using Primo have documented a number of tricks like this and &lt;a href=&#34;https://igelu.org/products-and-initiatives/product-working-groups/primo/special-projects/primo-community-support-primo-useful-bookmarklets/&#34;&gt;shared handy bookmarklets&lt;/a&gt; to rewrite urls and get catalogue data in different forms.&lt;/p&gt;
&lt;h2 id=&#34;iiif-and-images&#34;&gt;IIIF and images&lt;/h2&gt;
&lt;p&gt;SLV delivers digitised images using &lt;a href=&#34;https://iiif.io&#34;&gt;IIIF&lt;/a&gt;. The IIIF manifest urls are not directly exposed through the web interface, but you can construct your own.&lt;/p&gt;
&lt;p&gt;IIIF manifest urls look like this: &lt;a href=&#34;https://rosetta.slv.vic.gov.au/delivery/iiif/presentation/2.1/IE24074939/manifest.json&#34;&gt;https://rosetta.slv.vic.gov.au/delivery/iiif/presentation/2.1/IE24074939/manifest.json&lt;/a&gt; All we need to construct them is the &lt;code&gt;IE&lt;/code&gt; identifier, in this case &lt;code&gt;IE24074939&lt;/code&gt;. But where do you find this identifier?&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re looking at an image in the SLV&amp;rsquo;s image viewer, the url will be something like this: &lt;a href=&#34;https://viewer.slv.vic.gov.au/?entity=IE24074939&amp;amp;mode=browse&#34;&gt;https://viewer.slv.vic.gov.au/?entity=IE24074939&amp;amp;mode=browse&lt;/a&gt; Yep, the &lt;code&gt;IE&lt;/code&gt; identifier is right there in the url. Just extract it from the viewer url, and plug it into the manifest url!&lt;/p&gt;
&lt;p&gt;If you&amp;rsquo;re looking at a catalogue record, or starting with one of the &lt;code&gt;alma&lt;/code&gt; identifiers, you can get the &lt;code&gt;IE&lt;/code&gt; identifier from the &lt;code&gt;956$e&lt;/code&gt; field of the MARC record.&lt;/p&gt;
&lt;p&gt;The IIIF manifest will, in turn, provide identifiers for individual images that can be requested using the standard IIIF syntax.&lt;/p&gt;
&lt;p&gt;To save myself a bit of fiddling about, I created &lt;a href=&#34;https://gist.github.com/wragge/a37a4db854deffad956abc7bf918f6b0&#34;&gt;a userscript that exposes the IIIF manifest url&lt;/a&gt; within the image viewer. If you install it you&amp;rsquo;ll see something like this:&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-23-15-54-09.png&#34; width=&#34;510&#34; height=&#34;229&#34; alt=&#34;&#34;&gt;
&lt;h2 id=&#34;handles&#34;&gt;Handles&lt;/h2&gt;
&lt;p&gt;Links to digitised items sometimes come in the form of &amp;lsquo;handles&amp;rsquo;: &lt;a href=&#34;http://handle.slv.vic.gov.au/10381/4338980&#34;&gt;http://handle.slv.vic.gov.au/10381/4338980&lt;/a&gt; These urls are redirected to the image viewer.&lt;/p&gt;
&lt;p&gt;If you want to construct one of these handles, the identifier can be found in &lt;code&gt;956$a&lt;/code&gt; field of the MARC record.&lt;/p&gt;
&lt;h2 id=&#34;from-old-to-new&#34;&gt;From old to new&lt;/h2&gt;
&lt;p&gt;I was looking at the datasets created about 8 years ago in the &lt;a href=&#34;https://github.com/statelibraryvic/opendata&#34;&gt;SLV open data repository&lt;/a&gt; and noticed they included urls from the previous catalogue. Fortunately, the old urls redirect to the new system.&lt;/p&gt;
&lt;p&gt;For example, this url: &lt;a href=&#34;http://search.slv.vic.gov.au/MAIN:Everything:SLV_VOYAGER1842440&#34;&gt;http://search.slv.vic.gov.au/MAIN:Everything:SLV_VOYAGER1842440&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Redirects to: &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/fulldisplay?context=L&amp;amp;vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;docid=alma9918424403607636&#34;&gt;https://find.slv.vic.gov.au/discovery/fulldisplay?context=L&amp;amp;vid=61SLV_INST:SLV&amp;amp;search_scope=slv_local&amp;amp;tab=searchProfile&amp;amp;docid=alma9918424403607636&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;If you look closely at the urls you&amp;rsquo;ll see that the identifier from the old system is embedded in the new identifier: &lt;code&gt;1842440&lt;/code&gt; is in &lt;code&gt;9918424403607636&lt;/code&gt; – &lt;code&gt;99_1842440_3607636&lt;/code&gt;. This means if you have a lot of old urls, such as in the open datasets, you can easily rewrite them in your code.&lt;/p&gt;
&lt;h2 id=&#34;the-process-of-glam-hacking&#34;&gt;The process of GLAM hacking&lt;/h2&gt;
&lt;p&gt;No doubt a lot of this is well-known to librarians, and there&amp;rsquo;s probably many subtleties or complexities that my poking about has missed. But I wanted to document the process as much as the results – to give an idea of what I do when I approach a new GLAM collection online. I suppose this is GLAM hacking 101.&lt;/p&gt;
</description>
      <source:markdown>I like urls. They take you places. And if you know how to read them, they can tell you things about the systems that created them.
One of the first things I did when I started [my residency at SLV LAB](https://updates.timsherratt.org/2025/09/22/creative-technologistinresidence-at-the-state.html), was to try and understand how their collection urls work. There&#39;s a couple of well-worn methods I use when digging into a new site.

The first is url hacking – this involves fiddling around with the parameters in a url and submitting the result to see what happens. The Trove Data Guide includes [some examples of hacking Trove urls](https://tdg.glam-workbench.net/understanding-search/search-hacks.html) to change the delivery of search results.

The second method involves opening up the developer console in your web browser and watching the activity in the network tab as you click on links. This tells you where the information that gets loaded into your browser actually comes from – sometimes exposing handy urls that you can use to shortcut access to useful data.

## Permalinks

The SLV uses Primo for its public-facing catalogue, as well as other systems such as Rosetta and IIIF to deliver digitised content. I&#39;d noticed that [Zotero](https://www.zotero.org) gets some useful data from the catalogue using the default &#39;Primo 2018&#39; translator, however, important things like the item url aren&#39;t captured. The problem is that Primo&#39;s &#39;permalinks&#39; are generated as required by a browser click – they&#39;re not embedded anywhere on the page. This makes it hard to Zotero to grab them. So I started wondering how Zotero could construct short, persistent(ish) links to items.

Here&#39;s a link to an item in Primo: https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;search_scope=slv_local&amp;tab=searchProfile&amp;context=L&amp;docid=alma9941325055707636 

It looks pretty long and messy, but if you start deleting parameters and resubmitting, you&#39;ll find that only two parameters are essential, `vid` and `docid`. This means we can rewrite the url as: https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;docid=alma9941325055707636 Much nicer.

The &#39;permalink&#39; for the same item is: https://find.slv.vic.gov.au/permalink/61SLV_INST/1sev8ar/alma9941325055707636 If you look closely at the url path and compare it to the example above you&#39;ll see the path is constructed from `/vid/[some other id]/docid`. One of the librarians explained to me that the other identifier in the permalink is an encoding of the view type, but given that the &#39;fulldisplay&#39; view is the default, we don&#39;t really need it. So the shortened url seems fine for use in Zotero and is easy to generate from the current url. Nice.

It&#39;s also worth noting that the `vid` value doesn&#39;t seem to change, so to construct catalogue urls in your code, all you really need is the ALMA identifier that&#39;s in the `docid` parameter.

## Structured data

Item pages in Primo include a link labelled &#39;Display source record&#39;. If you click on this you&#39;re taken to a representation of the item&#39;s metadata in MARC. Here&#39;s what the urls look like: https://find.slv.vic.gov.au/discovery/sourceRecord?vid=61SLV_INST%3ASLV&amp;docId=alma9941325055707636&amp;recordOwner=61SLV_INST Notice that the &#39;fulldisplay&#39; in the url path above has changed to &#39;sourceRecord&#39;. There&#39;s also a new `recordOwner` parameter, but it seems you can delete this and still get the same result.

Having access to the MARC record is handy, because it delivers the metadata in a simple, structured plain text format. But while the &#39;source record&#39; page looks like a plain text file, it&#39;s actually a HTML page that embeds a plain text record. If you open up the network tab of your browser&#39;s developer console and reload the &#39;source record&#39; page, you&#39;ll see a different url is loaded under the hood: https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;vid=61SLV_INST:SLV&amp;recordOwner=61SLV_INST&amp;lang=en See how the url path has changed from `/discovery/` to `/primaws/rest/pub`? This url *does* deliver a plain text version of the MARC record.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-23-15-21-05.png&#34; width=&#34;600&#34; height=&#34;215&#34; alt=&#34;&#34;&gt;

Once you have the plain text version you can parse the contents to extract the structured data. There are tools that can probably do this automatically, but it&#39;s also pretty easy using regular expressions. Here&#39;s an example of some code I used to parse map records.

``` python
def get_marc_value(marc, tag, subfield):
    &#34;&#34;&#34;
    Gets the value of a tag/subfield from a text version of an item&#39;s MARC record.
    &#34;&#34;&#34;
    try:
        tag = re.search(rf&#34;^{tag}\t.+&#34;, marc, re.M).group(0)
        subfield = re.search(rf&#34;\${subfield}([^\$]+)&#34;, tag).group(1)
    except AttributeError:
        return None
    return subfield.strip(&#34; .,&#34;)
```
You can also access a JSON representation of the record by adding the parameter `&amp;showPnx=true` to the catalogue url: https://find.slv.vic.gov.au/discovery/fulldisplay?vid=61SLV_INST:SLV&amp;search_scope=slv_local&amp;tab=searchProfile&amp;context=L&amp;docid=alma9941325055707636&amp;showPnx=true

Once again, this is a JSON representation embedded in a web page. Using the same developer console trick, you can identify the direct url is: https://find.slv.vic.gov.au/primaws/rest/pub/pnxs/L/alma9941325055707636?vid=61SLV_INST:SLV&amp;lang=en&amp;search_scope=slv_local&amp;showPnx=true&amp;lang=en You should be able to parse the response from this url as JSON and use it in your code. I think the Zotero translator makes use of this `pnx` data.

If you want to download the MARC or JSON representations in your code, all you really need is the `alma` identifier. Just use it to construct one of the direct urls, such as this: https://find.slv.vic.gov.au/primaws/rest/pub/sourceRecord?docId=alma9941325055707636&amp;vid=61SLV_INST:SLV The `recordOwner` and `lang` parameters are not needed, and the `vid` parameter doesn&#39;t change.

Librarians using Primo have documented a number of tricks like this and [shared handy bookmarklets](https://igelu.org/products-and-initiatives/product-working-groups/primo/special-projects/primo-community-support-primo-useful-bookmarklets/) to rewrite urls and get catalogue data in different forms.

## IIIF and images

SLV delivers digitised images using [IIIF](https://iiif.io). The IIIF manifest urls are not directly exposed through the web interface, but you can construct your own.

IIIF manifest urls look like this: https://rosetta.slv.vic.gov.au/delivery/iiif/presentation/2.1/IE24074939/manifest.json All we need to construct them is the `IE` identifier, in this case `IE24074939`. But where do you find this identifier?

If you&#39;re looking at an image in the SLV&#39;s image viewer, the url will be something like this: https://viewer.slv.vic.gov.au/?entity=IE24074939&amp;mode=browse Yep, the `IE` identifier is right there in the url. Just extract it from the viewer url, and plug it into the manifest url!

If you&#39;re looking at a catalogue record, or starting with one of the `alma` identifiers, you can get the `IE` identifier from the `956$e` field of the MARC record.

The IIIF manifest will, in turn, provide identifiers for individual images that can be requested using the standard IIIF syntax.

To save myself a bit of fiddling about, I created [a userscript that exposes the IIIF manifest url](https://gist.github.com/wragge/a37a4db854deffad956abc7bf918f6b0) within the image viewer. If you install it you&#39;ll see something like this:

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-23-15-54-09.png&#34; width=&#34;510&#34; height=&#34;229&#34; alt=&#34;&#34;&gt;

## Handles

Links to digitised items sometimes come in the form of &#39;handles&#39;: http://handle.slv.vic.gov.au/10381/4338980 These urls are redirected to the image viewer.

If you want to construct one of these handles, the identifier can be found in `956$a` field of the MARC record.

## From old to new

I was looking at the datasets created about 8 years ago in the [SLV open data repository](https://github.com/statelibraryvic/opendata) and noticed they included urls from the previous catalogue. Fortunately, the old urls redirect to the new system.

For example, this url: http://search.slv.vic.gov.au/MAIN:Everything:SLV_VOYAGER1842440

Redirects to: https://find.slv.vic.gov.au/discovery/fulldisplay?context=L&amp;vid=61SLV_INST:SLV&amp;search_scope=slv_local&amp;tab=searchProfile&amp;docid=alma9918424403607636

If you look closely at the urls you&#39;ll see that the identifier from the old system is embedded in the new identifier: `1842440` is in `9918424403607636` – `99_1842440_3607636`. This means if you have a lot of old urls, such as in the open datasets, you can easily rewrite them in your code.

## The process of GLAM hacking

No doubt a lot of this is well-known to librarians, and there&#39;s probably many subtleties or complexities that my poking about has missed. But I wanted to document the process as much as the results – to give an idea of what I do when I approach a new GLAM collection online. I suppose this is GLAM hacking 101.







</source:markdown>
    </item>
    
    <item>
      <title>Creative Technologist-in-Residence at the State Library of Victoria!</title>
      <link>https://updates.timsherratt.org/2025/09/22/creative-technologistinresidence-at-the-state.html</link>
      <pubDate>Mon, 22 Sep 2025 23:14:09 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/09/22/creative-technologistinresidence-at-the-state.html</guid>
      <description>&lt;p&gt;I&amp;rsquo;m very excited to be the new &lt;a href=&#34;https://lab.slv.vic.gov.au/residencies-opportunities&#34;&gt;Creative Technologist-in-Residence at the SLV LAB&lt;/a&gt;. For the next few months I get to play around with metadata and images, think about online access, experiment with different technologies, and build things to help people to explore the State Library&amp;rsquo;s collections. In other words, I get to be in my happy place!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/2025-09-22-11.36.59.jpg&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;&#34;&gt;
&lt;p&gt;My group at &lt;a href=&#34;https://updates.timsherratt.org/2025/08/29/wikifest-at-the-state-library.html&#34;&gt;the recent SLV WikiFest&lt;/a&gt; was thinking about ways of helping researchers find resources relating to particular locations – how do I find material about my suburb, or my street? Coincidentally, the main focus of my residency will also be place-based collections, so I get to really think through some of the possibilities. SLV staff have already pointed me to some amazing maps and photographs, such as the &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;amp;collectionId=81271917420007636&#34;&gt;Committee for Urban Action collection&lt;/a&gt;, the &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/search?query=any,contains,mahlstedt%20melbourne&amp;amp;tab=searchProfile&amp;amp;search_scope=slv_local&amp;amp;vid=61SLV_INST:SLV&amp;amp;facet=tlevel,include,online_resources&amp;amp;offset=0&#34;&gt;Mahlstedt fire survey maps&lt;/a&gt;, the &lt;a href=&#34;https://guides.slv.vic.gov.au/MMBWplans&#34;&gt;MMBW plans&lt;/a&gt;, and the &lt;a href=&#34;https://find.slv.vic.gov.au/discovery/search?query=series,exact,Parish%20maps%20of%20Victoria&amp;amp;vid=61SLV_INST:SLV&amp;amp;offset=0&#34;&gt;Victorian parish maps&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;At the same time, I&amp;rsquo;ll be using my usual GLAM hacking approach to poke around in the SLV website to try and understand what data is currently available, identify any roadblocks, and document opportunities for computational research.&lt;/p&gt;
&lt;p&gt;The results of my residency will be shared on the &lt;a href=&#34;https://lab.slv.vic.gov.au&#34;&gt;SLV LAB site&lt;/a&gt;, in &lt;a href=&#34;https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency&#34;&gt;GitHub&lt;/a&gt;, in the &lt;a href=&#34;https://glam-workbench.net/state-library-victoria/&#34;&gt;SLV section of the GLAM Workbench&lt;/a&gt;, and of course here. As usual, I&amp;rsquo;ll be working in the open, documenting things as I go along, so please join me on the journey!&lt;/p&gt;
&lt;p&gt;Although the residency was formally announced today, I&amp;rsquo;ve actually been working with SLV data for the last couple of weeks and I&amp;rsquo;ve already got a backlog of stuff I need to blog about. Here&amp;rsquo;s a taster – what happens when you generate bounding boxes for thousands of parish maps from the available metadata and throw them on a map…?&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-22-23-08-25.png&#34; width=&#34;600&#34; height=&#34;406&#34; alt=&#34;&#34;&gt;
</description>
      <source:markdown>I&#39;m very excited to be the new [Creative Technologist-in-Residence at the SLV LAB](https://lab.slv.vic.gov.au/residencies-opportunities). For the next few months I get to play around with metadata and images, think about online access, experiment with different technologies, and build things to help people to explore the State Library&#39;s collections. In other words, I get to be in my happy place!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/2025-09-22-11.36.59.jpg&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;&#34;&gt;

My group at [the recent SLV WikiFest](https://updates.timsherratt.org/2025/08/29/wikifest-at-the-state-library.html) was thinking about ways of helping researchers find resources relating to particular locations – how do I find material about my suburb, or my street? Coincidentally, the main focus of my residency will also be place-based collections, so I get to really think through some of the possibilities. SLV staff have already pointed me to some amazing maps and photographs, such as the [Committee for Urban Action collection](https://find.slv.vic.gov.au/discovery/collectionDiscovery?vid=61SLV_INST:SLV&amp;collectionId=81271917420007636), the [Mahlstedt fire survey maps](https://find.slv.vic.gov.au/discovery/search?query=any,contains,mahlstedt%20melbourne&amp;tab=searchProfile&amp;search_scope=slv_local&amp;vid=61SLV_INST:SLV&amp;facet=tlevel,include,online_resources&amp;offset=0), the [MMBW plans](https://guides.slv.vic.gov.au/MMBWplans), and the [Victorian parish maps](https://find.slv.vic.gov.au/discovery/search?query=series,exact,Parish%20maps%20of%20Victoria&amp;vid=61SLV_INST:SLV&amp;offset=0).

At the same time, I&#39;ll be using my usual GLAM hacking approach to poke around in the SLV website to try and understand what data is currently available, identify any roadblocks, and document opportunities for computational research.

The results of my residency will be shared on the [SLV LAB site](https://lab.slv.vic.gov.au), in [GitHub](https://github.com/StateLibraryVictoria-SLVLAB/geo-maps-residency), in the [SLV section of the GLAM Workbench](https://glam-workbench.net/state-library-victoria/), and of course here. As usual, I&#39;ll be working in the open, documenting things as I go along, so please join me on the journey!

Although the residency was formally announced today, I&#39;ve actually been working with SLV data for the last couple of weeks and I&#39;ve already got a backlog of stuff I need to blog about. Here&#39;s a taster – what happens when you generate bounding boxes for thousands of parish maps from the available metadata and throw them on a map…?

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-09-22-23-08-25.png&#34; width=&#34;600&#34; height=&#34;406&#34; alt=&#34;&#34;&gt;








</source:markdown>
    </item>
    
    <item>
      <title>WikiFest at the State Library of Victoria</title>
      <link>https://updates.timsherratt.org/2025/08/29/wikifest-at-the-state-library.html</link>
      <pubDate>Fri, 29 Aug 2025 15:07:26 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/08/29/wikifest-at-the-state-library.html</guid>
      <description>&lt;p&gt;This week I was lucky enough to participate in WikiFest at the State Library of Victoria. Organised by the &lt;a href=&#34;https://lab.slv.vic.gov.au&#34;&gt;State Library&amp;rsquo;s new innovation LAB&lt;/a&gt; and &lt;a href=&#34;https://wikimedia.org.au&#34;&gt;Wikimedia Australia&lt;/a&gt;, Wikifest was a hands-on, participant-led workshop focused on the possibilities of connecting SLV&amp;rsquo;s collections to (and through!) Wikidata.&lt;/p&gt;
&lt;p&gt;The day kicked off with a series of presentations demonstrating possible uses of Wikidata. I talked a bit about some of my recent GLAM/Wikidata experiments. My &lt;a href=&#34;https://slides.com/wragge/wikifest-slv-2025&#34;&gt;slides are online&lt;/a&gt; and contain plenty of links to code, demonstrations, and documentation. They&amp;rsquo;re openly-licensed, so feel free to take anything of use.&lt;/p&gt;
&lt;iframe src=&#34;https://slides.com/wragge/wikifest-slv-2025/embed&#34; width=&#34;100%&#34; height=&#34;500&#34; title=&#34;WikiFest SLV 2025&#34; scrolling=&#34;no&#34; frameborder=&#34;0&#34; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;
&lt;p&gt;The rest of the day was spent in groups, working on particular projects and learning more about Wikidata in the process. My group was looking at providing placed-based entry points to SLV collections, and spent a lot of time exploring the representation of Victoria&amp;rsquo;s &lt;a href=&#34;https://query.wikidata.org/embed.html#%23Country%20populations%20together%20with%20total%20city%20populations%0ASELECT%20%3Flga%20%3FlgaLabel%20%3FstartDate%20%3FendDate%20%3Fpoint%20%7B%0A%20%20%3Flga%20wdt%3AP31%20wd%3AQ30129411%20%3B%0A%20%20%20%20%20%20%20wdt%3AP131%20wd%3AQ36687.%0A%20%20%3Flga%20p%3AP625%20%3Fcoordinate.%0A%20%20%3Fcoordinate%20ps%3AP625%20%3Fpoint.%0A%20%20OPTIONAL%20%7B%3Flga%20wdt%3AP571%20%3FstartDate.%7D%0A%20%20OPTIONAL%20%7B%3Flga%20wdt%3AP576%20%3FendDate.%7D%0A%20%20SERVICE%20wikibase%3Alabel%20%7B%20bd%3AserviceParam%20wikibase%3Alanguage%20%22%5BAUTO_LANGUAGE%5D%2Cmul%2Cen%22%20%7D%0A%7D&#34;&gt;Local Government Areas (LGAs) in Wikidata&lt;/a&gt;. We realised there was quite a bit of work to do in adding things like dates and boundaries, but we could see some exciting future possibilities. We also made a start, adding an &amp;lsquo;inception&amp;rsquo; date for the &lt;a href=&#34;https://www.wikidata.org/wiki/Q5123821&#34;&gt;City of Moe&lt;/a&gt;, based on the Victorian Government Gazette, &lt;a href=&#34;https://gazette.slv.vic.gov.au&#34;&gt;digitised by the SLV&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-08-29-14-45-09.png&#34; width=&#34;600&#34; height=&#34;271&#34; alt=&#34;Screen capture from Wikidata showing the inception property for the City of Moe&#34;&gt;
&lt;h2 id=&#34;bonus-userscript&#34;&gt;Bonus userscript&lt;/h2&gt;
&lt;p&gt;While I was preparing my presentation I was thinking about the the way entries for Australian people in Wikidata are linked to a range of different identifiers, such as DAAO, the Encyclopedia of Australian Science, and the Australian Dictionary of Biography (ADB). Often a single person can have multiple identifiers and this means that those identifiers themselves become connected through that person&amp;rsquo;s record. You can query Wikidata with one identifier, and get back links to a range of other information sources about that person.&lt;/p&gt;
&lt;p&gt;To demonstrate this, I created &lt;a href=&#34;https://gist.github.com/wragge/40f66af72c400b2563f95bda60e713dd&#34;&gt;a simple userscript&lt;/a&gt; that adds additional links to biographies in the ADB. The script grabs the ADB identifier from the url, queries Wikidata for additional identifiers, and writes the results into the page&amp;rsquo;s &amp;lsquo;Life Summary&amp;rsquo;. Basic, but useful!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/adb-userscript-example.png&#34; width=&#34;600&#34; height=&#34;253&#34; alt=&#34;Screenshot from the ADB showing the related links from Wikidata added to the Life Summary of Margaret Baskerville&#34;&gt;
&lt;p&gt;For something more advanced, have a look at the &lt;a href=&#34;https://addons.mozilla.org/en-US/firefox/addon/entity-explosion/&#34;&gt;Entity Explosion extension&lt;/a&gt; for Firefox.&lt;/p&gt;
</description>
      <source:markdown>This week I was lucky enough to participate in WikiFest at the State Library of Victoria. Organised by the [State Library&#39;s new innovation LAB](https://lab.slv.vic.gov.au) and [Wikimedia Australia](https://wikimedia.org.au), Wikifest was a hands-on, participant-led workshop focused on the possibilities of connecting SLV&#39;s collections to (and through!) Wikidata.

The day kicked off with a series of presentations demonstrating possible uses of Wikidata. I talked a bit about some of my recent GLAM/Wikidata experiments. My [slides are online](https://slides.com/wragge/wikifest-slv-2025) and contain plenty of links to code, demonstrations, and documentation. They&#39;re openly-licensed, so feel free to take anything of use.

&lt;iframe src=&#34;https://slides.com/wragge/wikifest-slv-2025/embed&#34; width=&#34;100%&#34; height=&#34;500&#34; title=&#34;WikiFest SLV 2025&#34; scrolling=&#34;no&#34; frameborder=&#34;0&#34; webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;&lt;/iframe&gt;

The rest of the day was spent in groups, working on particular projects and learning more about Wikidata in the process. My group was looking at providing placed-based entry points to SLV collections, and spent a lot of time exploring the representation of Victoria&#39;s [Local Government Areas (LGAs) in Wikidata](https://query.wikidata.org/embed.html#%23Country%20populations%20together%20with%20total%20city%20populations%0ASELECT%20%3Flga%20%3FlgaLabel%20%3FstartDate%20%3FendDate%20%3Fpoint%20%7B%0A%20%20%3Flga%20wdt%3AP31%20wd%3AQ30129411%20%3B%0A%20%20%20%20%20%20%20wdt%3AP131%20wd%3AQ36687.%0A%20%20%3Flga%20p%3AP625%20%3Fcoordinate.%0A%20%20%3Fcoordinate%20ps%3AP625%20%3Fpoint.%0A%20%20OPTIONAL%20%7B%3Flga%20wdt%3AP571%20%3FstartDate.%7D%0A%20%20OPTIONAL%20%7B%3Flga%20wdt%3AP576%20%3FendDate.%7D%0A%20%20SERVICE%20wikibase%3Alabel%20%7B%20bd%3AserviceParam%20wikibase%3Alanguage%20%22%5BAUTO_LANGUAGE%5D%2Cmul%2Cen%22%20%7D%0A%7D). We realised there was quite a bit of work to do in adding things like dates and boundaries, but we could see some exciting future possibilities. We also made a start, adding an &#39;inception&#39; date for the [City of Moe](https://www.wikidata.org/wiki/Q5123821), based on the Victorian Government Gazette, [digitised by the SLV](https://gazette.slv.vic.gov.au).

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-08-29-14-45-09.png&#34; width=&#34;600&#34; height=&#34;271&#34; alt=&#34;Screen capture from Wikidata showing the inception property for the City of Moe&#34;&gt;

## Bonus userscript

While I was preparing my presentation I was thinking about the the way entries for Australian people in Wikidata are linked to a range of different identifiers, such as DAAO, the Encyclopedia of Australian Science, and the Australian Dictionary of Biography (ADB). Often a single person can have multiple identifiers and this means that those identifiers themselves become connected through that person&#39;s record. You can query Wikidata with one identifier, and get back links to a range of other information sources about that person.

To demonstrate this, I created [a simple userscript](https://gist.github.com/wragge/40f66af72c400b2563f95bda60e713dd) that adds additional links to biographies in the ADB. The script grabs the ADB identifier from the url, queries Wikidata for additional identifiers, and writes the results into the page&#39;s &#39;Life Summary&#39;. Basic, but useful!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/adb-userscript-example.png&#34; width=&#34;600&#34; height=&#34;253&#34; alt=&#34;Screenshot from the ADB showing the related links from Wikidata added to the Life Summary of Margaret Baskerville&#34;&gt;

For something more advanced, have a look at the [Entity Explosion extension](https://addons.mozilla.org/en-US/firefox/addon/entity-explosion/) for Firefox.


</source:markdown>
    </item>
    
    <item>
      <title>GLAM hacking with userscripts</title>
      <link>https://updates.timsherratt.org/2025/07/17/glam-hacking-with-userscripts.html</link>
      <pubDate>Thu, 17 Jul 2025 17:21:25 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/07/17/glam-hacking-with-userscripts.html</guid>
      <description>&lt;p&gt;In teaching and workshops I used to get students to question the idea that websites are &amp;lsquo;published&amp;rsquo;. They&amp;rsquo;re not released into the world in a fixed, immutable form – they&amp;rsquo;re a set of blueprints which only reach their final form in your browser window. This makes it possible to change the way websites look and behave.&lt;/p&gt;
&lt;p&gt;Mozilla used to have a nifty educational tool called X-Ray Googles. Using it you could explore the code underlying a web page and do fun things like inserting new text or images. I encouraged students to try hacking ASIO&amp;rsquo;s home page.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/asio-eggplants.jpg&#34; width=&#34;600&#34; height=&#34;629&#34; alt=&#34;Old, modified screenshot of ASIO homepage with a section of a cartoon from First Dog On the Moon inserted.&#34;&gt;
&lt;p&gt;&lt;em&gt;ASIO home page with some added First Dog on the Moon.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;There are other ways you can fiddle with websites. For example, most browsers have a developer console that exposes the code and styling of a page. You can use the console to edit HTML elements or toggle styles, but your changes won&amp;rsquo;t be saved.&lt;/p&gt;
&lt;p&gt;One way you can save and share your web site customisations is by creating userscripts. Userscripts are little bits of Javascript code that run in your browser after a web page loads. These scripts can change many aspects of a page – not just how it looks, but also how it works.&lt;/p&gt;
&lt;h2 id=&#34;some-old-userscripts&#34;&gt;Some old userscripts&lt;/h2&gt;
&lt;p&gt;I&amp;rsquo;ve been playing around with userscripts for a long time. Back in 2008, I created a userscript that &lt;a href=&#34;https://discontents.com.au/shoebox/archives-shoebox/archives-in-3d.html&#34;&gt;completely overhauled the way that digital files were presented&lt;/a&gt; in the National Archives of Australia&amp;rsquo;s online database, RecordSearch. My userscript added new options for navigating and printing the file, and even made it possible to view the complete file contents on a 3D zoomable wall.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/userscript-screenshot1.jpg&#34; width=&#34;600&#34; height=&#34;577&#34; alt=&#34;Screenshot of a digitised file in RecordSearch showing the features added by the userscript.&#34;&gt;
&lt;p&gt;&lt;em&gt;This customised RecordSearch interface was created by a userscript.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The userscripts I&amp;rsquo;ve created over the years have tended to either be useful little hacks aimed at fixing annoying aspects of GLAM websites, or experiments in thinking about the sort of information that&amp;rsquo;s presented online by GLAM organisations, and how it might be different.&lt;/p&gt;
&lt;p&gt;In the first category are hacks like my &lt;a href=&#34;https://gist.github.com/wragge/b2af9dc56f7cb0a9476b&#34;&gt;RecordSearch show pages userscript&lt;/a&gt;. In 2009, I got annoyed that there was no way of knowing how many pages were in a digitised file until you clicked on the link. &lt;a href=&#34;https://discontents.com.au/doing-it-yourself/index.html&#34;&gt;So I fixed it.&lt;/a&gt; With my userscript running, the links to digitised files are rewritten to display the number of pages. I&amp;rsquo;ve updated the code numerous times over the years, adding new features, and dealing with changes to RecordSearch. The last update was just a few days ago.&lt;/p&gt;
&lt;p&gt;In the second category is my userscript that inserts photos from &lt;a href=&#34;https://www.realfaceofwhiteaustralia.net/&#34;&gt;The Real Face of White Australia&lt;/a&gt; into RecordSearch. The are many thousands of records in the National Archives of Australia that document the impact of the White Australia Policy on the lives of ordinary people. But it&amp;rsquo;s often hard to understand this from the file descriptions. The userscript displays portrait images extracted from the files alongside the metadata – it tells you there are &lt;a href=&#34;https://doi.org/10.5281/zenodo.3579530&#34;&gt;people inside&lt;/a&gt;.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/people-inside-list.gif&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;Animated gif showing how the userscript changes the display of a list of files in RecordSearch by adding pictures of people.&#34;&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/people-inside-item.gif&#34; width=&#34;600&#34; height=&#34;412&#34; alt=&#34;Animated gif showing how the userscript changes the display of an individual files in RecordSearch by adding pictures of the people inside.&#34;&gt;
&lt;p&gt;Amidst &lt;a href=&#34;https://updates.timsherratt.org/2025/07/09/the-rebirth-of-wragge-labs.html&#34;&gt;my recent self-archiving binge&lt;/a&gt;, I realised I&amp;rsquo;d never updated this userscript to work with the latest data from The Real Face of White Australia, so I spent some time getting it working again. In the process I realised that RecordSearch now included content security policies that made it a bit harder to insert new images. The solution was to use one of the special userscript functions, &lt;a href=&#34;https://www.tampermonkey.net/documentation.php?locale=en#api:GM_addElement&#34;&gt;GM_addElement()&lt;/a&gt;, rather than plain old Javascript. But then I discovered that the if the show pages userscript ran after this one, it would trigger the security restrictions nonetheless! To avoid this I made sure that the two userscripts operated on separate elements. So now the &lt;a href=&#34;https://gist.github.com/wragge/2941e473ee70152f4de7&#34;&gt;show people userscript&lt;/a&gt; is working again!&lt;/p&gt;
&lt;h2 id=&#34;and-a-new-userscript-to-improve-trove-lists&#34;&gt;And a new userscript to improve Trove lists&lt;/h2&gt;
&lt;p&gt;Fixing up the &amp;lsquo;people inside&amp;rsquo; code reminded me of how much fun it was playing around with userscripts, so when David Coombe mentioned a problem he had using Trove lists on Mastodon last night, I had to have a go at fixing it.&lt;/p&gt;
&lt;p&gt;The problem is that Trove lists display all the tags associated with each individual item. Some items have lots of tags, so this eats up the screen real estate, making it harder to browse the contents of a list. Notes attached to items can be hidden, but not tags. Why not?&lt;/p&gt;
&lt;p&gt;My &lt;a href=&#34;https://gist.github.com/wragge/ab6a9d6b612bee6bc4d98658e947c420&#34;&gt;brand new userscript&lt;/a&gt; hides tags by default, and adds a new link to toggle their visibility for each individual item. The link also displays the number of tags attached to each item. This gives the user control over which tags are displayed and when.&lt;/p&gt;
&lt;p&gt;&lt;video src=&#34;https://cdn.uploads.micro.blog/8371/2025/simplescreenrecorder-2025-07-17-12.53.06.mp4&#34; poster=&#34;https://updates.timsherratt.org/uploads/2025/poster.png&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34; width=&#34;600px&#34;&gt;&lt;/video&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;The new userscript in action – toggle your tags!&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;The main difficulty in creating this userscript was knowing when the page had actually finished loading. The current version of Trove uses a lot of Javascript to load and manipulate content, so you have to tell the userscript to wait until  everything has settled down. Otherwise the script could fire too soon and cause unexpected results. I tried a number of different approaches to handling this problem, but eventually settled on the &lt;a href=&#34;https://gist.github.com/BrockA/2625891&#34;&gt;waitForKeyElements script&lt;/a&gt;. (I just realised there&amp;rsquo;s a &lt;a href=&#34;https://github.com/CoeJoder/waitForKeyElements.js&#34;&gt;more recent version&lt;/a&gt; of this script that doesn&amp;rsquo;t require JQuery, so I might need to investigate this further.)&lt;/p&gt;
&lt;p&gt;Another Trove problem fixed!&lt;/p&gt;
&lt;h2 id=&#34;using-userscripts&#34;&gt;Using userscripts&lt;/h2&gt;
&lt;p&gt;In addition to the userscripts mentioned above, I&amp;rsquo;ve also created one that &lt;a href=&#34;https://gist.github.com/wragge/af8bd20a14005d267ffc759463bd832c&#34;&gt;enables you to browse Trove newspaper pages using the arrows on your keyboard&lt;/a&gt;. Left and right arrows go to the next and previous pages, while up and down arrows jump between issues. Searching is great, but sometimes you just want to browse. Install this userscript for that old-time, authentic newspaper reading experience!&lt;/p&gt;
&lt;p&gt;But how do you install userscripts? First of all you need a browser extension to manage your userscripts – I use &lt;a href=&#34;https://www.tampermonkey.net/&#34;&gt;TamperMonkey&lt;/a&gt; or &lt;a href=&#34;http://violentmonkey.com/&#34;&gt;ViolentMonkey&lt;/a&gt;. Just follow the instructions to add one of them to your browser.&lt;/p&gt;
&lt;p&gt;To install one of my userscripts, you need to go to the script (saved as a GitHub Gist) and click on the &amp;lsquo;Raw&amp;rsquo; button. Your userscript manager will then ask you if you want to add the userscript. Click install!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-07-17-16-59-11.png&#34; width=&#34;600&#34; height=&#34;244&#34; alt=&#34;&#34;&gt;
&lt;p&gt;&lt;em&gt;Click on the &amp;lsquo;Raw&amp;rsquo; button to install.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Once they&amp;rsquo;re installed the userscripts will run automatically when specified pages are loaded. If you ever want to disable them, you can do that from your userscript manager&amp;rsquo;s dashboard.&lt;/p&gt;
&lt;p&gt;For convenience, here are the Gist links to all the userscripts I&amp;rsquo;ve mentioned:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href=&#34;https://gist.github.com/wragge/b2af9dc56f7cb0a9476b&#34;&gt;RecordSearch show pages&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://gist.github.com/wragge/2941e473ee70152f4de7&#34;&gt;RecordSearch show people&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://gist.github.com/wragge/ab6a9d6b612bee6bc4d98658e947c420&#34;&gt;Trove lists hide tags&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href=&#34;https://gist.github.com/wragge/af8bd20a14005d267ffc759463bd832c&#34;&gt;Trove newspapers keyboard navigation&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;As with anything you install on your computer, you want to make sure that you trust the source of any userscripts you add.&lt;/p&gt;
</description>
      <source:markdown>In teaching and workshops I used to get students to question the idea that websites are &#39;published&#39;. They&#39;re not released into the world in a fixed, immutable form – they&#39;re a set of blueprints which only reach their final form in your browser window. This makes it possible to change the way websites look and behave.

Mozilla used to have a nifty educational tool called X-Ray Googles. Using it you could explore the code underlying a web page and do fun things like inserting new text or images. I encouraged students to try hacking ASIO&#39;s home page.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/asio-eggplants.jpg&#34; width=&#34;600&#34; height=&#34;629&#34; alt=&#34;Old, modified screenshot of ASIO homepage with a section of a cartoon from First Dog On the Moon inserted.&#34;&gt;

*ASIO home page with some added First Dog on the Moon.*

There are other ways you can fiddle with websites. For example, most browsers have a developer console that exposes the code and styling of a page. You can use the console to edit HTML elements or toggle styles, but your changes won&#39;t be saved.

One way you can save and share your web site customisations is by creating userscripts. Userscripts are little bits of Javascript code that run in your browser after a web page loads. These scripts can change many aspects of a page – not just how it looks, but also how it works.

## Some old userscripts

I&#39;ve been playing around with userscripts for a long time. Back in 2008, I created a userscript that [completely overhauled the way that digital files were presented](https://discontents.com.au/shoebox/archives-shoebox/archives-in-3d.html) in the National Archives of Australia&#39;s online database, RecordSearch. My userscript added new options for navigating and printing the file, and even made it possible to view the complete file contents on a 3D zoomable wall.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/userscript-screenshot1.jpg&#34; width=&#34;600&#34; height=&#34;577&#34; alt=&#34;Screenshot of a digitised file in RecordSearch showing the features added by the userscript.&#34;&gt;

*This customised RecordSearch interface was created by a userscript.*

The userscripts I&#39;ve created over the years have tended to either be useful little hacks aimed at fixing annoying aspects of GLAM websites, or experiments in thinking about the sort of information that&#39;s presented online by GLAM organisations, and how it might be different.

In the first category are hacks like my [RecordSearch show pages userscript](https://gist.github.com/wragge/b2af9dc56f7cb0a9476b). In 2009, I got annoyed that there was no way of knowing how many pages were in a digitised file until you clicked on the link. [So I fixed it.](https://discontents.com.au/doing-it-yourself/index.html) With my userscript running, the links to digitised files are rewritten to display the number of pages. I&#39;ve updated the code numerous times over the years, adding new features, and dealing with changes to RecordSearch. The last update was just a few days ago.

In the second category is my userscript that inserts photos from [The Real Face of White Australia](https://www.realfaceofwhiteaustralia.net/) into RecordSearch. The are many thousands of records in the National Archives of Australia that document the impact of the White Australia Policy on the lives of ordinary people. But it&#39;s often hard to understand this from the file descriptions. The userscript displays portrait images extracted from the files alongside the metadata – it tells you there are [people inside](https://doi.org/10.5281/zenodo.3579530). 

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/people-inside-list.gif&#34; width=&#34;600&#34; height=&#34;450&#34; alt=&#34;Animated gif showing how the userscript changes the display of a list of files in RecordSearch by adding pictures of people.&#34;&gt;

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/people-inside-item.gif&#34; width=&#34;600&#34; height=&#34;412&#34; alt=&#34;Animated gif showing how the userscript changes the display of an individual files in RecordSearch by adding pictures of the people inside.&#34;&gt;

Amidst [my recent self-archiving binge](https://updates.timsherratt.org/2025/07/09/the-rebirth-of-wragge-labs.html), I realised I&#39;d never updated this userscript to work with the latest data from The Real Face of White Australia, so I spent some time getting it working again. In the process I realised that RecordSearch now included content security policies that made it a bit harder to insert new images. The solution was to use one of the special userscript functions, [GM_addElement()](https://www.tampermonkey.net/documentation.php?locale=en#api:GM_addElement), rather than plain old Javascript. But then I discovered that the if the show pages userscript ran after this one, it would trigger the security restrictions nonetheless! To avoid this I made sure that the two userscripts operated on separate elements. So now the [show people userscript](https://gist.github.com/wragge/2941e473ee70152f4de7) is working again!

## And a new userscript to improve Trove lists

Fixing up the &#39;people inside&#39; code reminded me of how much fun it was playing around with userscripts, so when David Coombe mentioned a problem he had using Trove lists on Mastodon last night, I had to have a go at fixing it.

The problem is that Trove lists display all the tags associated with each individual item. Some items have lots of tags, so this eats up the screen real estate, making it harder to browse the contents of a list. Notes attached to items can be hidden, but not tags. Why not?

My [brand new userscript](https://gist.github.com/wragge/ab6a9d6b612bee6bc4d98658e947c420) hides tags by default, and adds a new link to toggle their visibility for each individual item. The link also displays the number of tags attached to each item. This gives the user control over which tags are displayed and when.

&lt;video src=&#34;https://cdn.uploads.micro.blog/8371/2025/simplescreenrecorder-2025-07-17-12.53.06.mp4&#34; poster=&#34;https://updates.timsherratt.org/uploads/2025/poster.png&#34; controls=&#34;controls&#34; preload=&#34;metadata&#34; width=&#34;600px&#34;&gt;&lt;/video&gt;

*The new userscript in action – toggle your tags!*

The main difficulty in creating this userscript was knowing when the page had actually finished loading. The current version of Trove uses a lot of Javascript to load and manipulate content, so you have to tell the userscript to wait until  everything has settled down. Otherwise the script could fire too soon and cause unexpected results. I tried a number of different approaches to handling this problem, but eventually settled on the [waitForKeyElements script](https://gist.github.com/BrockA/2625891). (I just realised there&#39;s a [more recent version](https://github.com/CoeJoder/waitForKeyElements.js) of this script that doesn&#39;t require JQuery, so I might need to investigate this further.)

Another Trove problem fixed!

## Using userscripts

In addition to the userscripts mentioned above, I&#39;ve also created one that [enables you to browse Trove newspaper pages using the arrows on your keyboard](https://gist.github.com/wragge/af8bd20a14005d267ffc759463bd832c). Left and right arrows go to the next and previous pages, while up and down arrows jump between issues. Searching is great, but sometimes you just want to browse. Install this userscript for that old-time, authentic newspaper reading experience!

But how do you install userscripts? First of all you need a browser extension to manage your userscripts – I use [TamperMonkey](https://www.tampermonkey.net/) or [ViolentMonkey](http://violentmonkey.com/). Just follow the instructions to add one of them to your browser.

To install one of my userscripts, you need to go to the script (saved as a GitHub Gist) and click on the &#39;Raw&#39; button. Your userscript manager will then ask you if you want to add the userscript. Click install!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/screenshot-from-2025-07-17-16-59-11.png&#34; width=&#34;600&#34; height=&#34;244&#34; alt=&#34;&#34;&gt;

*Click on the &#39;Raw&#39; button to install.*

Once they&#39;re installed the userscripts will run automatically when specified pages are loaded. If you ever want to disable them, you can do that from your userscript manager&#39;s dashboard.

For convenience, here are the Gist links to all the userscripts I&#39;ve mentioned:

- [RecordSearch show pages](https://gist.github.com/wragge/b2af9dc56f7cb0a9476b)
- [RecordSearch show people](https://gist.github.com/wragge/2941e473ee70152f4de7)
- [Trove lists hide tags](https://gist.github.com/wragge/ab6a9d6b612bee6bc4d98658e947c420)
- [Trove newspapers keyboard navigation](https://gist.github.com/wragge/af8bd20a14005d267ffc759463bd832c)

As with anything you install on your computer, you want to make sure that you trust the source of any userscripts you add.
</source:markdown>
    </item>
    
    <item>
      <title>The rebirth of Wragge Labs (and moving my Heroku apps)</title>
      <link>https://updates.timsherratt.org/2025/07/09/the-rebirth-of-wragge-labs.html</link>
      <pubDate>Wed, 09 Jul 2025 16:48:23 +1000</pubDate>
      
      <guid>http://wragge.micro.blog/2025/07/09/the-rebirth-of-wragge-labs.html</guid>
      <description>&lt;p&gt;It looks like some paid work I was counting on won&amp;rsquo;t be going ahead, so I&amp;rsquo;m trying to save a bit of money on cloud hosting. As I previously noted, this resulted in &lt;a href=&#34;https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html&#34;&gt;the resurrection of &lt;em&gt;The future of the past&lt;/em&gt;&lt;/a&gt;, but I&amp;rsquo;ve also been continuing to slog away at migrating all my old Flask apps and experiments from Heroku to a single Digital Ocean droplet. As of today, I&amp;rsquo;ve migrated 11 apps. Here&amp;rsquo;s a few details&amp;hellip;&lt;/p&gt;
&lt;h2 id=&#34;a-new-old-home&#34;&gt;A new (old) home&lt;/h2&gt;
&lt;p&gt;The first thing I had to figure out was how to group together a series of individual &lt;a href=&#34;https://flask.palletsprojects.com/en/stable/&#34;&gt;Flask&lt;/a&gt; apps so I could easily run and maintain them on a single server, without making major changes to the apps themselves. I decided to go with the &lt;a href=&#34;https://flask.palletsprojects.com/en/stable/patterns/appdispatch/&#34;&gt;application dispatching pattern&lt;/a&gt; described in the Flask documentation. This groups the apps within a single Python environment so I had to do some alignment of Python versions and packages, but it wasn&amp;rsquo;t too hard and having just one virtual environment to manage seems a lot easier in the long run.&lt;/p&gt;
&lt;p&gt;The application dispatching pattern configures the server to run one application at the web root (&#39;/&#39;), with the other apps assigned individual sub-paths. This raised the question, what did I want sitting at the root address? Rather than selecting an existing application for the prime slot, I decided to take the opportunity to build a showcase that included details of many of the things I&amp;rsquo;ve created over the years.&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/wraggelabs.png&#34; width=&#34;600&#34; height=&#34;720&#34; alt=&#34;Screenshot of the original Wragge Labs&#34;&gt;
&lt;p&gt;&lt;em&gt;The old Wragge Labs (circa 2012)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;I also needed a new domain name. Or did I? Back in the old days, I had a site where I shared many of my tools and experiments – Wragge Labs. In the intervening years, I&amp;rsquo;d moved or migrated much of the content away and pointed the wraggelabs.com domain to my main site at timsherratt.au. But this seemed like a good opportunity to resurrect it. So if you&amp;rsquo;d like to have a play around with some of the things I&amp;rsquo;ve created over the last 30 years, head along to the all new &lt;a href=&#34;https://wraggelabs.com&#34;&gt;wraggelabs.com&lt;/a&gt;!&lt;/p&gt;
&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/new-wraggelabs.png&#34; width=&#34;600&#34; height=&#34;569&#34; alt=&#34;Screenshot of part of the new Wragge Labs!&#34;&gt;
&lt;p&gt;&lt;em&gt;The new &lt;a href=&#34;https://wraggelabs.com&#34;&gt;Wragge Labs&lt;/a&gt; showcases websites, apps, and experiments from the past 30 years – some useful, some playful, and some creepy&amp;hellip;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;A surprising amount of the things I&amp;rsquo;ve built are still working. As I was compiling my list, I made a few running repairs – for example, fixing broken links in &lt;a href=&#34;https://timsherratt.au/shed/culturevic/&#34;&gt;Linking history in place&lt;/a&gt; and &lt;a href=&#34;https://timsherratt.au/shed/magicsquares/&#34;&gt;Magic Squares&lt;/a&gt; to get them working again. However, some things only exist now in web archives, and others have been broken by the recent actions of the &lt;a href=&#34;https://updates.timsherratt.org/2025/05/07/farewell-trove.html&#34;&gt;NLA&lt;/a&gt; and &lt;a href=&#34;https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html&#34;&gt;NAA&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;To provide a bit of extra context, I&amp;rsquo;ve grouped together publications and presentations documenting many of the experiments. These are saved in &lt;a href=&#34;https://www.zotero.org/&#34;&gt;Zotero&lt;/a&gt;, and tagged with the name of the app. When you click on a &amp;lsquo;Related&amp;rsquo; link, the details of any linked resources are retrieved using the Zotero API and displayed on a new page. This means I can add new related resources simply by dropping  them into Zotero.&lt;/p&gt;
&lt;h2 id=&#34;moving-house&#34;&gt;Moving house&lt;/h2&gt;
&lt;p&gt;The process for moving the apps to their new home was pretty straightforward. On my local machine, I copied the code into the new aggregated structure, added any packages needed into a combined requirements file, and created a new top-level app to direct requests. And then I spun everything up and started fixing bugs&amp;hellip;&lt;/p&gt;
&lt;p&gt;All of the problems were easily resolved. Most involved fixing up paths to static assets, or in navigation links. The only significant changes to the Python code were caused by the deprecation of the &lt;code&gt;.count()&lt;/code&gt; method in Pymongo.&lt;/p&gt;
&lt;p&gt;To make life a little harder, I decided to take the opportunity to make sure that all the assets – javascript, css, and font files – were loaded from the local system, and not sitting in the cloud. Having everything local should make it easier to maintain the apps in the long term. It was a bit fiddly tracking down where everything was being loaded from, but not too hard.&lt;/p&gt;
&lt;p&gt;The only other changes I made were to add some caching to most of the apps, particularly those that make calls to external databases or APIs. I used &lt;a href=&#34;https://flask-caching.readthedocs.io/en/latest/index.html&#34;&gt;Flask-Caching&lt;/a&gt; with the local file system backend.&lt;/p&gt;
&lt;p&gt;To get the new aggregated application working on a Digital Ocean droplet, I followed the instructions on how to &lt;a href=&#34;https://www.digitalocean.com/community/tutorials/how-to-serve-flask-applications-with-uwsgi-and-nginx-on-ubuntu-22-04&#34;&gt;serve Flask applications using uWSGI and Nginx&lt;/a&gt;. I think the only thing I did differently was to use &lt;a href=&#34;https://github.com/pyenv/pyenv&#34;&gt;pyenv&lt;/a&gt; to manage Python versions and the virtual environment. To update the app, I use &lt;code&gt;rsync&lt;/code&gt; to copy across the code and &lt;code&gt;systemctl&lt;/code&gt; to restart it. So far it&amp;rsquo;s all working pretty smoothly.&lt;/p&gt;
&lt;h2 id=&#34;redirecting-heroku&#34;&gt;Redirecting Heroku&lt;/h2&gt;
&lt;p&gt;Once the apps were happy in their new home, I needed to redirect the Heroku addresses to &lt;a href=&#34;https://wraggelabs.com&#34;&gt;wraggelabs.com&lt;/a&gt;. It was surprisingly hard to find good documentation on how to do this, so I&amp;rsquo;ll document my steps in detail in case its of use to others.&lt;/p&gt;
&lt;p&gt;There are a few redirect apps for Heroku around, but I decided to use &lt;a href=&#34;https://github.com/fastmonkeys/heroku-redirect&#34;&gt;heroku-redirect&lt;/a&gt; because it basically just configures and runs Nginx without any additional processing. First I cloned &lt;code&gt;heroku-redirect&lt;/code&gt; to my local system, and then for each app I wanted to migrate I followed these steps:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;cd into the &lt;code&gt;heroku-redirect&lt;/code&gt; directory&lt;/li&gt;
&lt;li&gt;set the git remote for the app you want to redirect: &lt;code&gt;heroku git:remote -a [app name]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;if you&amp;rsquo;re not using the lastest Heroku stack, update it: &lt;code&gt;heroku stack:set heroku-24 -a [app name]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;set the url you want to redirect to (without trailing slash): &lt;code&gt;heroku config:set LOCATION=[new url]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;to add the full path to the redirected url: &lt;code&gt;heroku config:set PRESERVE_PATH=true&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;I found I had to remove the Python buildpack from the app before updating, I did this using the Heroku dashboard, but no doubt there&amp;rsquo;s also a CLI command&lt;/li&gt;
&lt;li&gt;I also used the dashboard to add a new nginx buildpack: &lt;code&gt;https://github.com/heroku/heroku-buildpack-nginx.git&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;finally I pushed the new app, using &lt;code&gt;--force&lt;/code&gt; to replace it completely: &lt;code&gt;git push --force heroku master:main&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Of course, you&amp;rsquo;ll need to have the Heroku CLI installed.&lt;/p&gt;
&lt;p&gt;A number of my Heroku apps were using basic dynos (which cost US$7 a month), so once they were redirected, I changed them to used shared eco dynos. Yay – money saved! Hopefully, the redirects won&amp;rsquo;t push the eco dynos beyond their monthly limit.&lt;/p&gt;
&lt;h2 id=&#34;more-experiments-to-come&#34;&gt;More experiments to come?&lt;/h2&gt;
&lt;p&gt;One of the good things about all of this housekeeping is that it&amp;rsquo;s got me thinking about new experiments. I used to love Flask and Heroku because it they made it so easy to build and share things. Now I can do the same with &lt;a href=&#34;https://wraggelabs.com&#34;&gt;wraggelabs.com&lt;/a&gt;!&lt;/p&gt;
</description>
      <source:markdown>It looks like some paid work I was counting on won&#39;t be going ahead, so I&#39;m trying to save a bit of money on cloud hosting. As I previously noted, this resulted in [the resurrection of *The future of the past*](https://updates.timsherratt.org/2025/07/02/the-future-of-the-past.html), but I&#39;ve also been continuing to slog away at migrating all my old Flask apps and experiments from Heroku to a single Digital Ocean droplet. As of today, I&#39;ve migrated 11 apps. Here&#39;s a few details...
## A new (old) home

The first thing I had to figure out was how to group together a series of individual [Flask](https://flask.palletsprojects.com/en/stable/) apps so I could easily run and maintain them on a single server, without making major changes to the apps themselves. I decided to go with the [application dispatching pattern](https://flask.palletsprojects.com/en/stable/patterns/appdispatch/) described in the Flask documentation. This groups the apps within a single Python environment so I had to do some alignment of Python versions and packages, but it wasn&#39;t too hard and having just one virtual environment to manage seems a lot easier in the long run.

The application dispatching pattern configures the server to run one application at the web root (&#39;/&#39;), with the other apps assigned individual sub-paths. This raised the question, what did I want sitting at the root address? Rather than selecting an existing application for the prime slot, I decided to take the opportunity to build a showcase that included details of many of the things I&#39;ve created over the years.

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/wraggelabs.png&#34; width=&#34;600&#34; height=&#34;720&#34; alt=&#34;Screenshot of the original Wragge Labs&#34;&gt;

*The old Wragge Labs (circa 2012)*

I also needed a new domain name. Or did I? Back in the old days, I had a site where I shared many of my tools and experiments – Wragge Labs. In the intervening years, I&#39;d moved or migrated much of the content away and pointed the wraggelabs.com domain to my main site at timsherratt.au. But this seemed like a good opportunity to resurrect it. So if you&#39;d like to have a play around with some of the things I&#39;ve created over the last 30 years, head along to the all new [wraggelabs.com](https://wraggelabs.com)!

&lt;img src=&#34;https://cdn.uploads.micro.blog/8371/2025/new-wraggelabs.png&#34; width=&#34;600&#34; height=&#34;569&#34; alt=&#34;Screenshot of part of the new Wragge Labs!&#34;&gt;

*The new [Wragge Labs](https://wraggelabs.com) showcases websites, apps, and experiments from the past 30 years – some useful, some playful, and some creepy...*

A surprising amount of the things I&#39;ve built are still working. As I was compiling my list, I made a few running repairs – for example, fixing broken links in [Linking history in place](https://timsherratt.au/shed/culturevic/) and [Magic Squares](https://timsherratt.au/shed/magicsquares/) to get them working again. However, some things only exist now in web archives, and others have been broken by the recent actions of the [NLA](https://updates.timsherratt.org/2025/05/07/farewell-trove.html) and [NAA](https://updates.timsherratt.org/2025/05/19/no-more-harvesting-data-from.html).

To provide a bit of extra context, I&#39;ve grouped together publications and presentations documenting many of the experiments. These are saved in [Zotero](https://www.zotero.org/), and tagged with the name of the app. When you click on a &#39;Related&#39; link, the details of any linked resources are retrieved using the Zotero API and displayed on a new page. This means I can add new related resources simply by dropping  them into Zotero.

## Moving house

The process for moving the apps to their new home was pretty straightforward. On my local machine, I copied the code into the new aggregated structure, added any packages needed into a combined requirements file, and created a new top-level app to direct requests. And then I spun everything up and started fixing bugs...

All of the problems were easily resolved. Most involved fixing up paths to static assets, or in navigation links. The only significant changes to the Python code were caused by the deprecation of the `.count()` method in Pymongo.

To make life a little harder, I decided to take the opportunity to make sure that all the assets – javascript, css, and font files – were loaded from the local system, and not sitting in the cloud. Having everything local should make it easier to maintain the apps in the long term. It was a bit fiddly tracking down where everything was being loaded from, but not too hard.

The only other changes I made were to add some caching to most of the apps, particularly those that make calls to external databases or APIs. I used [Flask-Caching](https://flask-caching.readthedocs.io/en/latest/index.html) with the local file system backend.

To get the new aggregated application working on a Digital Ocean droplet, I followed the instructions on how to [serve Flask applications using uWSGI and Nginx](https://www.digitalocean.com/community/tutorials/how-to-serve-flask-applications-with-uwsgi-and-nginx-on-ubuntu-22-04). I think the only thing I did differently was to use [pyenv](https://github.com/pyenv/pyenv) to manage Python versions and the virtual environment. To update the app, I use `rsync` to copy across the code and `systemctl` to restart it. So far it&#39;s all working pretty smoothly.

## Redirecting Heroku

Once the apps were happy in their new home, I needed to redirect the Heroku addresses to [wraggelabs.com](https://wraggelabs.com). It was surprisingly hard to find good documentation on how to do this, so I&#39;ll document my steps in detail in case its of use to others.

There are a few redirect apps for Heroku around, but I decided to use [heroku-redirect](https://github.com/fastmonkeys/heroku-redirect) because it basically just configures and runs Nginx without any additional processing. First I cloned `heroku-redirect` to my local system, and then for each app I wanted to migrate I followed these steps:

- cd into the `heroku-redirect` directory
- set the git remote for the app you want to redirect: `heroku git:remote -a [app name]`
- if you&#39;re not using the lastest Heroku stack, update it: `heroku stack:set heroku-24 -a [app name]`
- set the url you want to redirect to (without trailing slash): `heroku config:set LOCATION=[new url]`
- to add the full path to the redirected url: `heroku config:set PRESERVE_PATH=true`
- I found I had to remove the Python buildpack from the app before updating, I did this using the Heroku dashboard, but no doubt there&#39;s also a CLI command
- I also used the dashboard to add a new nginx buildpack: `https://github.com/heroku/heroku-buildpack-nginx.git`
- finally I pushed the new app, using `--force` to replace it completely: `git push --force heroku master:main`

Of course, you&#39;ll need to have the Heroku CLI installed.

A number of my Heroku apps were using basic dynos (which cost US$7 a month), so once they were redirected, I changed them to used shared eco dynos. Yay – money saved! Hopefully, the redirects won&#39;t push the eco dynos beyond their monthly limit.

## More experiments to come?

One of the good things about all of this housekeeping is that it&#39;s got me thinking about new experiments. I used to love Flask and Heroku because it they made it so easy to build and share things. Now I can do the same with [wraggelabs.com](https://wraggelabs.com)! 









</source:markdown>
    </item>
    
  </channel>
</rss>
