Webarchiv: 20 Years of Web Archiving in the Czech Republic

By Marie Haškovcová, Illyria Brejchová, Luboš Svoboda, and Andrea Prokopová (Czech Web Archive of National Library of the Czech Republic)

An Introduction to Webarchiv

The idea to create a national web archive that would preserve the growing amount of Czech digital born media was conceived as soon as 1999. In the year 2000, Webarchiv was founded as a joint project of the National Library of the Czech Republic, the Moravian Library, and the Masaryk University, making it one of the oldest webarchives. The first websites were archived in 2001, regular harvesting began in 2005, and in 2007 Webarchiv joined the IIPC.

Webarchiv home page

Currently, Webarchiv is part of the National Library of the Czech Republic and holds approx. 400 TB of data. Webarchive collects this data in a variety of ways. Through comprehensive harvests, second order domains of *.cz are harvested once or twice a year thanks to a cooperation with the Czech domain provider CZ.NIC (currently it is about 1.4 million URLs). Czech web resources with historical, scientific or cultural value are selectively harvested more frequently and in depth compared to comprehensive harvests. Finally, resources connected to a specific event or topical are collected through topic harvests. Webarchiv currently has more than 30 topical collections, covering elections, Olympics, climate change and more. Continuous harvesting (automated, several times a day) is currently being tested on some thematic collections, such as COVID-19 or Czech media.

A gif showing the development of the Webarchiv website created using Time Map Visualization

 

Data harvesting and accessibility

The big challenge for Webarchiv at the moment is assuring the accessibility of the data in its collection, both in regards to maintaining the ability to display the archived websites, but also in regards to allowing access to its collection to researchers as well as the public. In terms of public access, only 0,4 % of the whole collection is available freely online. This is due to current Czech legislation which allows the National Library to make reproductions of a work for its own archiving and conservation purposes, but does not entitle libraries to make them available. Online access is therefore made available only to resources in the selective harvest which are licensed under a Creative Commons licence or after signing a contract with the publisher. Websites available to the public are catalogued in accordance with the RDA rules and integrated into the Czech national bibliography. The entire collection can be accessed by the public on the library premises.

Webarchiv catalogue

On the technical side, Webarchiv uses open source software, such as Heritrix 3.4 and OpenWayback 3.0, but also develops its own open source tools, such as Seeder for managing electronic resources, websites and harvests or WA-KAT as an online resource cataloguing tool. We are testing the harvesting of social media accounts of politicians, are experimenting with UMBRA, and apply manual harvesting using Webrecorder 2.3, which allows curators to harvest web 2.0 or more technologically complex websites such as online exhibitions or multimedia magazines. We plan to replace the current Wayback 3.0 application with Python Wayback (pywb) to display this type of content. We consider the autonomy of curators in regard to planning harvests and quality assurance to be key even in automated harvests, which is why we continue to improve Seeder, a tool for managing harvests and curating web resources. In the future, Seeder should allow curators to perform harvests without technical support, which will allow them to react more efficiently to the ephemeral online environment.

Seeder – tool for managing electronic resources, websites and harvests

 

Collaboration with key partners

Over the years, Webarchiv has developed a collaboration with various institutions. Notably, these include the aforementioned CZ.NIC, the Institute of Czech Literature of the Czech Academy of Sciences, for whom we are archiving the online Czech literary tradition from the beginning of the Czech Internet to the present day, or the Czech National Archive, for whom we are archiving websites of public agencies, such as ministries or other central administrative authorities. As for international collaboration, we worked with the University Library in Bratislava on a shared topic collection of online resources relating to the 30th anniversary of the Velvet Revolution, which led to the collapse of the communist regime in former Czechoslovakia, and regularly contribute to the IIPC collaborative collections, most recently to the COVID-19 collection.

Topical collections

 

As for making our collection more accessible to researches Webarchiv is involved in a research project titled “Development of a centralized interface for extracting big data from web archives”. The National Library of the Czech Republic partnered with the the Department of Cybernetics of the Faculty of Applied Sciences at the University of West Bohemia and the Institute of Sociology of the Czech Academy of Sciences on this project focused on Webarchiv´s data research. The main aim of the project is to develop a centralized user interface which would allow researchers to search through data collected by Webarchiv and obtain datasets for further research. The outcome of the project will be a faceted full text search engine for analyzing large quantities of web archive data with an integrated application for exporting selected datasets. The research project is expected to be completed in 2022.

Engaging the public

Website nomination form

 

Webarchiv also actively engages with the public. we accept suggestions for resources to include in selective harvests on our website, and are also active on social media. We have a long-going campaign, where we share dead websites from our collection (websites that no longer exist, but we have archive copies) on Facebook and we have recently started doing the same on Instagram. Through these activities, we hope to raise awareness and interest in web archiving. We are also active on Twitter, where we recently participated in the #WarcnetChallenge.

A contribution of Webarchiv to the Warcnet Challenge

We also launched a new blog, where we post a series called 10 websites for eternity, where personalities from various fields share a list of Czech websites they can not imagine life without and which they would regret losing if they were discontinued without being preserved in an archive. It is an opportunity for them to share their top 10 treasures the Czech web has to offer, both forgotten webs worm their bookmarks, and accomplished veterans of the internet. Webarchiv then archives the websites on the list and adds them to a topic collection. We see opening Webarchiv to curatorship from external specialists from various fields as an important direction to head in and a great way to expand our current curatorship strategies. Another topic we are considering is Link Rot, we see it as an area in which we can prove to be very beneficial – in the future, a catalog of valuable web resources could be created, curated directly by scientists and students.

From series 10 websites for eternity

The future of Webarchiv

Similar to other web archives, Webarchiv is facing numerous challenges – not only in the field of acquisition, preservation and access to data or legislative changes. The internet is an ever changing environment and we are therefore always a step behind in our efforts to preserve it. Numerous questions offer themselves up for debate: How should we approach ethical questions regarding the use of data collected during harvests? How can content on social media be preserved within its appropriate context when feeds are personalized? Should we preserve software along with the archived web pages? How should we approach valuable online content accessible only beyond a paywall? We are excited to be part of Webarchiv s journey and witness the ways in which it tackles these questions, and matures in the following decades!

The Danish Coronavirus web collection – Coronavirus on the curators’ minds

By Sabine Schostag, Web Curator, The Royal Danish Library

Introduction – a provoking cartoon

In a sense, the story of Corona and the national Danish Web Archive (Netarchive) starts at the end of January 2020 – about 6 weeks before Corona came to Denmark. A cartoon by Niels Bo Bojesens in the Danish newspaper “Jyllandsposten” (2020-01-26) showing the Chinese flag with a circle of yellow corona-viruses instead of the stars caused indignation in China and captured attention worldwide. We focused on collecting reactions on different social media and in the international news media. Particularly on Twitter, a seething discussion arose with vehement comments and memes about Denmark.

From epidemic to pandemic

After that, the curators again focused on the daily routines in web archiving, as we believed that Corona (Covid-19) was a closed chapter in Netarchive’s history. But this was not the case. When the IIPC Content Development Working Group launched the Covid-19 collection in February, the Royal Danish Library contributed the Danish seeds.

Suddenly, the Corona virus arrived in Europe and the first infected Dane came home from a skiing trip in Italy. The epidemic turned into a pandemic. On March 12, the Danish Government decided to lockdown the country: all public employees where sent to their home offices and borders were closed. Not only the public sector shut down, trade and industry, shops, restaurants, bars etc. had to close too. Only supermarkets were still open and people in the Health Care sector had to work overtime.

While Denmark came to a standstill, so to speak, the Netarchive curators worked at full throttle on the coronavirus event collection. Zoom became the most important work tool for the following 2½ months. In daily Zoom meetings, we coordinated who worked on which facet of this collection. To put it briefly, we curators had coronavirus on our minds.

Event crawls in Netarchive

The Danish Web Archive crawls all Danish news media between several times daily and one time weekly, so there is no need to include news articles in an event crawl. Thus, with an event crawl we focus on augmented activity on social media, blog articles, new sites emerging in connection to the event – and reactions in news media outside Denmark.

Coronavirus documentation in Denmark

The Danish Web collection on coronavirus in Denmark is part of a general documentation on the corona lockdown in Denmark in 2020. This documentation is a cooperation between several cultural institutions, the National Archives (Rigsarkivet), the National Museum (Nationalmuseet), the Workers Museum (Arbejdermuseet), local archives and, last but not least, the Royal Danish Library. The corona lockdown documentation was supposed to be done in two steps:  the “here and now” collection of documentation under the corona lockdown and a more systematic follow-up by collecting materials from authorities and public bodies.

“Days with Corona” – a call for help

All Danes were asked to contribute to the corona lockdown documentation, for instance by sending photos and narratives from their daily life under the lockdown. “Days with Corona” is the title of this part of the documentation of the Danish Folklore Archives run by the National Museum and the Royal Library.

Netarchive also asked the public for help by nominating URLs of web pages related to coronavirus, social media profiles, hashtags, memes and any other relevant material.

Help from colleagues

Web archiving is part of the Department for Digital Cultural Heritage at the Royal Library. Almost all colleagues from the department were able to continue with their every day work from their home offices. Many colleagues from other departments were not able to do so. Some of them helped the Netarchive team by nominating URLs, as this event crawl could keep curators busy more than 7½ hours a day. We used a Google spreadsheet for all nominations (fig. 1)

Fig. 1 Nomination sheet for curators and colleagues form other departments and a call for contributions.

The Queen’s 80th birthday

On April 16, Queen Margarethe II celebrated her 80th birthday. One of the first things she did after the Corona lockdown, on March 13, was to cancel all her birthday celebration events. In a way, she set a good example, as everybody was asked not to meet with no more than ten people, ideally we only should socialize with members of our own household.

As part of the Corona event crawl, we collected web activity related to the Queen’s birthday, which mainly consisted of reactions on social media.

The big challenge – capturing social media

Knowledge of the coronavirus Covid-19 changes continuously. Consequently, authorities, public bodies, private institutions, and companies change information and precaution rules on their webpages frequently. We try to capture as much of these changes as possible. Companies and private individuals offering safety gear for protection against the virus was another facet in the collection. However, capturing all relevant activity on social media was much more challenging than the frequent updates on traditional web pages. Most of the social media platforms use technologies, which Heritrix (used by Netarchive for event crawling) is not able to capture.

Fig. 2 The Queen’s speech to the Danes on how to cope with the corona crisis. This was the second time in history (the first time was during the World War II) when a Royal Head of State addressed  the nation, besides the annual New Year’s Eve speech.

More or less successfully, we tried to capture content from Facebook, TikTok, Twitter, YouTube, Instagram, Reddit, Imgur, Soundcloud, and Pinterest. Twitter is the platform we are able to crawl with Heritrix with rather good results. We collect Facebook profiles with an account at Archive-It, as they have a better set of tools for capturing Facebook. With frequent Quality Assurance and follow-ups, we also get rather good results from Instagram, TikTok and Reddit. We capture YouTube videos by crawling the watch-URLs with a specific configuration using YouTube dl.  One of the collected YouTube videos comes from the Royal family’s YouTube channel: the Queens address to the people on how to behave to prevent or limit the spreading of the coronavirus (https://www.youtube.com/watch?v=TZKVUQ-E-UI, Fig. 2).

As Heritrix has problems with dynamic web content and streaming, we also used Webrecorder.io, although we have not yet implemented this tool in our harvesting setup. However, captures with Webrecorder.io are only drops in the ocean. The use of Webrecorder.io is manual: a curator clicks on all the elements on a page we want to capture. An example is a page on the BBC website, with a video of the reopening of Danish primary schools after the total lockdown (https://www.bbc.com/news/av/world-europe-52649919/coronavirus-inside-a-reopened-primary-school-in-the-time-of-covid-19, Fig. 3). There is still an issue with ingesting the resulting WARC files from Webrecorder.io in our web archive.

Danes produced a range of podcasts on coronavirus issues. We crawled the podcasts we had identified. We get good results when having an URL to a RSS feed, which we crawl with XML extraction.

Fig. 3 Crawled with Webrecorder.io to get the video.

Capture as much as possible – a broad crawl

Netarchive runs up to four broad crawls a year. We launched our first broad crawl for 2020 just in the beginning of the Danish Corona lockdown – on March 14. A broad crawl is an in-depth snapshot of all dk-domains and all other Top Level Domains (TDLs) where we have identified Danish content. A side benefit of this broad crawl might be getting Corona-related content into the archive – content which the curators do not find with their different methods. We identify content both with classic/common? keyword searches and using a variety of link scraping tools / link scrapers.

Is the coronavirus related web collection of any value to anybody?

In accordance with the Danish personal data protection law, the public has no access to the archived web material. Only researchers affiliated with Danish research institutions can apply for access in connection with specific research projects. We have already received an application for one research project dealing with values in the Covid-19 communication. We hope that our collection will inspire more research projects.

The Croatian Web Archive – what’s new?

The Croatian Web Archive (Hrvatski arhiv weba, HAW), launched in 2004, is open access. To celebrate its 15th anniversary, the National and University Library in Zagreb hosted the IIPC General Assembly and the Web Archiving Conference in June 2019. HAW has been the central point in Croatia for researching website development (.hr domain) and the HAW Team has also been organising training for librarians. One of HAW’s most recent projects was the development of the new portal.


By Karolina Holub, Library Adviser at the Croatian Digital Library Development Centre, Croatian Institute for Librarianship, Ingeborg Rudomino, Senior Librarian at the Croatian Web Archive, & Marta Matijević, Librarian at the Croatian Web Archive (National and University Library in Zagreb)

June 2019 – June 2020

It’s been more than a year since the National and University Library in Zagreb (NSK) hosted the IIPC General Assembly and Web Archiving Conference, which we remember with nostalgia.

Last year was a very busy year for the Croatian Web Archive (HAW) and we would like to share some of the key projects that we have been working on.

New portal design

The highlight of the last period was the launch of the new HAW portal.

Croatian Web Archive (HAW)

It was a complex project that took two years – from the initial idea to the launch of the portal in February 2020. The portal was developed and is maintained by NSK website developers and the HAW team. It is developed in a customized WordPress theme. Since the new portal had to be integrated with the database of the archived content, that is maintained by our partner University of Zagreb University Computing Centre (SRCE), a lot of coding was required in order to connect the portal with the archive database to ensure that everything is working properly and smoothly.

Below you can see fractions of our previous portals from 2006 and from 2020:

HAW’s website from 2006 until 2011

HAW’s website from 2011 until 2020

So, what’s new?

The most important objective was to put search box in focus for all types of crawls and give users an easier way to find a resource. Because of the diverse ways of searching, our goal was to have a clear distinction between selective (that is indexed and can be searched by keywords, any word in title or URL, or use advanced search) and domain crawls (can only be searched by entering the full URL). A valuable addition to this version of the portal are the basic metadata elements that accompany each resource (which has a catalogue record) available in the portal.


Archived resource with the basic metadata elements (available also via library catalogue)

Additionally, the browsing of subject categories has been expanded with subject subcategories.

The visibility of the thematic collections has been improved by placing them on the title page. A new feature In Focus has also been added to highlight some of the most important or interesting events or anniversaries happening in the country, city or at the Library in the form of blog posts. This feature is available only in the Croatian version of the portal. The central part of the homepage features New in HAW and Gone from the web sections where user can browse all publications that are new or publications that are no longer available on the live web. The About HAW page features a timeline marking all the important dates related to history of HAW.

Some parts of the new portal have largely remained the same with only slight improvements to make them more user-friendly and up to date. More information about Selection criteria, National .hr domain crawls, Statistics, Bibliography, FAQ etc. can be found in the footer.

The portal is also available in English.

New thematic collections

During this one-year period, we have been working on six thematic collections. Some of them are already available and others are still ongoing:

Elections for the President of the Republic of Croatia 2019-2020

At the end of 2019, Presidential Elections were held in Croatia. The thematic crawls was conducted in January and the content is publicly available as part of this thematic collection.

Rijeka – European Capital of Culture 2020

Croatian city of Rijeka is European Capital of Culture 2020. All contents related to this event, during this challenging time, will be harvested. We are still collecting the content.

Croatian Presidency of the Council of the European Union

Croatia has chaired the Council of the European Union from January to June 2020. We are finishing this thematic collection and it will soon be publicly available on the HAW’s portal.

COVID-19

Our largest thematic collection so far is definitely COVID-19, which is still ongoing. We have included the public in collecting the content inviting nominations related to the coronavirus. In this thematic collection, we follow the events that begin with the onset of coronavirus in the Republic of Croatia and the world, featured on the Croatian portals, blogs, articles – from the outbreak of coronavirus, through general lockdown to the gradual normalization in which we are now.

Archived website (19.03.2020)

2020 Zagreb earthquake

On March 22, just a few days after the start of coronavirus lockdown in Croatia, Zagreb was hit by the biggest earthquake in 140 years, causing numerous injuries and extensive damage. Croatian Web Archive immediately started collecting content about this disaster. This thematic collection is publicly available on the HAW’s portal.

Archived website (15.04.2020) (photo by HINA; Damir Senčar)

2020 Parliamentary Elections

When the spread of the coronavirus was believed to be under control, Croatia held the Parliamentary Elections on July 5. The content for this collection will be collected until the constitution of the new Croatian Parliament.

In May of this year, we started cataloguing thematic collections at the collection level. We have also contributed the Croatian content to the IIPC Coronavirus (Covid-19) Collection.

Annual .hr crawl

In December 2019 we have conducted the 9th annual domain crawl and collected 119 million resources amounting to 9.3 TB.

HAW also started the installation and configuration of tools for indexing and enabling full-text search for domain and thematic crawls: Webarchive-Discovery for parsing and indexing WARC files, Apache SORL for indexing and searching text content and SHINE web interface for index search and analysis. We are still in the testing phase and only a part of existing crawled content is indexed.

Testing Web Curator Tool for new collaborative processes – Local Web Crowd crawls

A new development phase is the collaboration with public libraries in crawling their local history collections for which we are testing the Web Curator Tool. We expect the first results are by the end of November this year.

What’s next?

In the next months, we will be working on enabling more advanced use of HAW’s content to better suit the researchers, starting with the creation of the data sets from HAW collections. We will also prepare guidelines for using archived content on HAW’s portal. In addition, we are planning to update our training material according to the new IIPC training material. In the meantime, we invite you to explore our new portal.

Documenting COVID-19 and the Great Confinement in Canada

By Sylvain Bélanger, Director General, Transition Team, Library and Archives Canada and Treasurer, International Internet Preservation Consortium

It seemed like it happened overnight, suddenly we were told to work from home and limit our physical interactions with people outside our household until further notice. The information was changing and evolving very rapidly and as we started seeing the rise in COVID-19 related cases globally, the anxiety among colleagues and employees was rising as well. Business rapidly ground to an almost complete halt and only essential services would continue to operate, with strict controls and restrictions.

Spanish Flu and the Great Confinement of 2020

Even during these early days, in these times of uncertainty, a group of individuals saw a parallel between the current situation and the period of the Spanish Flu a century earlier. Thinking ahead to fifty years from now this group was asking the question – how will future generations know about this period of time, the Great Confinement of 2020 as they may call it, or the time of great creativity, or perhaps the time the Internet became our lifeline? Turning the clock back one hundred years to the period of the Spanish Flu has given us hints. Let’s not forget that the tragedy of the early 1900s was documented through newspapers, diaries, photographs, and publications detailing the fight and aftermaths of the Spanish Flu.

In 2020, where social media and websites are key means citizens used to document and get informed, how do we capture such ephemeral product?  Does any country have the answer? Isn’t that the question we often ask ourselves?

The importance of web archiving

Screenshot from the Public Health Agency of Canada website.

This period has given all of us an opportunity to educate news publishers, citizens, and government decisions makers about the work done by web archiving teams across Canada and around the world. The efforts of the IIPC have been pushed to the forefront in this crisis, and have helped us demonstrate the importance of preserving web content for future generations.

In Canada the work entails a coordination of efforts with other governmental institutions as well as with university libraries and provincial/territorial archives to limit duplication of efforts. At Library and Archives Canada (LAC), to ensure a proper reflection of Canadian society, we have captured over 662,000 Tweets with hashtags such as #covidcanada, #covid19canada, #canadalockdown, #canadacovid19, as part of over 38 million digital assets collected for COVID-19 in 2020. Of that a little over 87% of the content is non-governmental, from media and non-media web resources selected for the COVID-19 collection. This includes 33 sites on Canadian news and media collected daily, to ensure we capture a robust sample of the published news on COVID-19. Added to that are non-media web resources that create an overall LAC seed list of over 900 resources. Total data collected to date is a little more than 3.09 TB at LAC alone.

Documenting the Canadian response

In addition to our web archiving program, LAC librarians have noticed an increase in books being published about the crisis. That has been measured through our ISBN team observing an increase in authors requesting ISBN numbers for books about various aspects of the pandemic. In addition, LAC will document the Government of Canada’s response to the COVID-19 pandemic through our Government Records Disposition Program.  In this way the government decision-making on COVID-19 and impact on Canadians will be acquired and preserved by LAC for present and future generations. Also, our Private Archives personnel are monitoring the activities, responses and reactions of individuals, communities, organizations and associations within their respective portfolios. LAC will endeavour to acquire documents about the pandemic when discussing possible acquisitions with current and potential donors and when evaluating offers. Descriptions in archival fonds will now highlight COVID 19 content where appropriate.

The efforts undertaken to date at LAC are meant to document the Canadian response. Are our efforts enough to help citizens 100 years from now to understand the times we were living, and how we responded to and tackled the challenges of COVID-19? Only time will tell whether this is enough, or we need to do any more work to truly document the historical times we live in.

From pilot to portal: a year of web archiving in Hungary

National Széchényi Library started a web archiving pilot project in 2017. The aim of the pilot project was to identify the requirements of establishing the Hungarian Internet Archive. In the two years of the pilot phase, some hundred cultural and scientific websites were selected and published with the owners’ permission. The Hungarian Web Archive (MIA) was officially launched in 2017. The Library joined the IIPC in 2018 and the Hungarian Web Archive was first introduced at the General Assembly in Wellington in 2018. Last year, the achievements of the project were presented at the Web Archiving Conference (WAC) in Zagreb, in June 2019. This blog post offers a summary of some key developments since the 2019 conference.


By Márton Németh, Digital librarian at the National Széchényi Library, Hungary

In just about a year, we moved from a pilot project to officially launching our web archive, running a comprehensive crawl and creating special collections. In May 2020, the Hungarian parliament passed the modifications of the Cultural Law which allows us to run web archiving activities as a part of its basic service portfolio. Over the past year we have also organised training and participated in various collaborative initiatives.

Conferences and collaborations

In the summer just after the Zagreb conference, we could exchange experiences with our Czech and Slovak colleagues about the current status and major development points of web archiving projects in the Czech Republic, Slovakia and Hungary in the Visegrad 4 Library Conference in Bratislava. Our presentation is available from here. In the autumn, at the annual international conference of digital preservation in Bratislava, we could elaborate on our basic thoughts about the potential use of microdata in library environment. The presentation can be downloaded from here.

At the Digital Humanities 2020 conference in Budapest, Hungary, we organized a whole web archiving session with presentations and panel discussions together with Marie Haskovcová from the Czech National Library, Kees Teszelszky from the National Library of the Netherlands, Balázs Indig from the Digital Humanities Research Centre of Loránd Eötvös University and with Márton Németh from the National Széchényi Library. The main aim was to get a spotlight on Digital Humanities research activities in the web archiving context. Our presentation is available from here.

Training

Our annual workshop in the National Széchényi Library focused on the metadata enrichment of web archives, crawling and managing local web content in university library and city library environments, crawling and managing online newspaper articles and setting the limits of web archiving in research library environments.

We also run several accredited training courses for Hungarian librarians and summarized our experiences in web archiving education field in an article published by Emerald. The membership in the IIPC Training Working Group has offered us valuable experiences in this field.

Domain crawl and new portal

We had run our second comprehensive harvest about a large segment of the Hungarian web domain in the end of 2019. The robot had started on 246.819 seed addresses and crawled 110 million URL-s in less than eight days with 6,4 TB storage.

Our original project website was the first repository of resources related to web archiving in Hungarian. In 2019 we built a new portal. This new website serves as a knowledgebase in web archiving field in Hungary. Beyond the introduction to the web archive and to the project, separate groups of resources (info-materials, documents etc.) are available for every-day users, for content-owners, for professional experts and for journalists. It is available at https://webarchivum.oszk.hu.

https://webarchivum.oszk.hu
webarchivum.oszk.hu

We created a new sub-collection in 2019-2020 on the Francis II Rákóczi Memorial Year at the National Széchényi Library (NSZL), within the framework of the Public Collection Digitization Strategy. Its primary goal was present the technology of web archiving and the integration of the web archive with other digital collections through a demo application. The content focuses on the webpages and websites related to the Memorial Year, to the War of Independence, to the Prince and to his family. Furthermore, it contains born digital or digitized books from the Hungarian Electronic Library, articles from the Electronic Periodical Archives, photos, illustrations and other visual documents from the Digital Archive of Pictures. The service is available on the following address: http://rakoczi2019.webarchivum.oszk.hu.

OSZK-figure2
rakoczi2019.webarchivum.oszk.hu

Legislation and new collections

In May 2020 the Hungarian parliament passed the modifications of the Cultural Law that entitles the National Széchényi Library to run web archiving activities as a part of its basic service portfolio. Legal deposit of web materials will also be established. The corresponding governmental and ministerial decrees will appear soon, all the law modifications and decrees will be in effect from 1 January 2021.

We made our first experiment of harvesting various materials from 700 pages with more than 100.000 posts from Instagram using the Webrecorder software. We are running event-based harvests too about COVID-19, Summer Olympic Games, Paris Peace Conference (1919-1920). We are joining also to the corresponding international IIPC collaborative collection development projects.

Next steps

Supported by the framework of the Public Collection Digitization Strategy we could start to develop a collaboration network with various regional libraries in Hungary in order to collect local materials for the Hungarian Web Archive. Hopefully, we will summarize our first experiences during our next annual workshop in the autumn and we can further develop our joint collection activities.

Luxembourg Web Archive – Coronavirus Response

By Ben Els, Digital Curator, The National Library of Luxembourg

The National Library of Luxembourg has been harvesting the Luxembourg web under the digital legal deposit since 2016. In addition to the large-scale domain crawls, the Luxembourg Web Archive also operates targeted crawls, aimed at specific subjects or events. During the past weeks and months, the global pandemic of the Coronavirus, has put society before unprecedented challenges. While large parts of our professional and social lives had to move even further online, the need to capture and document the implications of this crisis on the Internet, has seen enormous support in all domains of society. While it is safe to admit that web archiving is still a relatively unknown concept to most people in Luxembourg (probably also in other countries), it is also safe to say, that we have never seen a better case to illustrate the necessity of web archiving and ask for support in this overwhelming challenge.

webarchive.lu

Media and communities

At the National Library, we started our Coronavirus collection on March 16th, while there were 81 known cases in Luxembourg. While we have been harvesting websites in several event crawls for the past 3 years, it was clear from the start that the amount of information to be captured would surpass any other subject by a great deal. Therefore, we decided to ask for support from the Luxembourg news media, by asking them to send us lists of related news articles from their websites. This appeal to editors quickly evolved into a call for participation to the general public, asking all communities, associations, and civil interest groups to share their responses and online information about the crisis. Addressing the news media in the first place, gave us great support in spreading the word about the collection. Part of our approach to building an event collection, is to follow the news and take in information about new developments and publications of different organisations and persons of interest. As the flow and high-paced rhythm of new public information and support was vital to many communities, we also had to try and keep up with new websites, support groups and solidarity platforms being launched every day. However, many of these initiatives are not covered equally in the news or social media, a situation which is even more complicated through Luxembourg’s multilingual makeup. We learned about the challenges from the government and administrations, to convey important and urgent information in 4 or 5 languages at a time: Luxembourgish, French, German, English and Portuguese. The same goes for news and social media, and as a result, for the Luxembourg Web Archive. Therefore, we were grateful to receive contributions from organisations, which we would not have thought of including ourselves, and who were not talked about as much in the news.

© The Luxembourg Government

Effort and resources

While the need and support for web archiving exploded during March and April, it was also clear, that the standard resources allocated to the yearly operations of the web archive would not suffice in responding to the challenge in front of us. The National Library was able to increase our efforts, by securing additional funding, which allowed us to launch an impromptu domain crawl and to expand the data budget on Archive-It crawls. We are all aware of the uphill battle in communicating the benefits of archiving the web. There is a feeling that, while people generally agree on the necessity of preserving websites, in most cases there is little sense of urgency or immediate requirement – since after all, most everyday changes are perceived as corrections of mistakes, or improvements on previous versions. In my opinion, the case of Coronavirus related websites, made the idea of web archiving as a service and obligation to society much clearer and easier to convey.

© Ministry of Health

Private and public

The Web offers many spaces and facets for personal expression and communication. While social media have played a crucial part in helping people to deal with the crisis, web archives face some of their biggest challenges in harvesting and preserving social media. Alongside the technical difficulties and enormous related costs, there is the question of ethics in collecting content which is not 100% private, but also not 100% public. For instance, in Luxembourg, many support groups launched on Facebook, where people could ask their questions about the current situation and new developments in terms of what is

allowed, find help and comfort to their uncertainties. There are several active groups in every language, even some dedicated to districts of the city, with neighbours looking after each other. While it is important to try to capture all facets of an event (especially if this information is unique to the Internet) I am uncertain, whether it is ethical to capture the questions, comments and conversations of people in vulnerable situations. Even though there are sometimes thousands of members per group and pretty much everyone can join, they are not fully open to the public.

Collecting and sharing

covidmemory.lu

Besides the large-scale crawls and Archive-It collection, we also contributed part of our seed list to the IIPC’s collaborative Novel Coronavirus collection, led by the Content Development Working Group. Of course, the National Library did not limit its response to archiving websites. With our call for participation, we also received a variety of physical and digital documents: mainly from municipalities and public administrations who submitted numerous documents, which were issued to the public in relation the reorganisation of public services and the temporary restrictions on social life.

We also received some unexpected contributions, in the form of poems, essays and short diary entries written during confinement, describing and reflecting upon the current situation from a very personal angle. Likewise, a researcher shared his private bibliometric analysis of scientific literature about the Coronavirus. Furthermore, the University of Luxembourg’s Centre for Contemporary and Digital History has launched the sharing platform covidmemory.lu, enabling ordinary people living or working in Luxembourg to share their photos, videos, stories and interviews related to COVID-19.

Web Archiving Week 2021

Since the 2021 edition of the IIPC Web Archiving Conference will be part of the Web Archiving Week, in  partnership with the University of Luxembourg and the RESAW network, I am not going to spoil too much about the program by saying that we will continue exploring these shared efforts and responses during the week of June 14th – 18th 2021. We are looking forward to welcoming you all to Luxembourg!

Covid-19 Collecting at the National Library of New Zealand

By Gillian Lee, Coordinator, Web Archives at the Alexander Turnbull Library, National Library of New Zealand

The National Library of New Zealand reflects on their rapid response collecting of Covid-19 related websites since February 2020.

Collecting in response to the pandemic

Web Archivists at the National Library of New Zealand are used to collecting websites relating to major events, but the Covid-19 pandemic has had such a global impact, it’s affected every member of society. It has been heart breaking to see the tragic loss of life and economic hardships that people are facing world-wide. The effects of this pandemic will be with us for a long time.

Collecting content relating to these events always produces mixed emotions as a web archivist. There’s the tension between collecting content before it disappears, and in that regard, we put on our hard hats and get on with it. At the same time however, these events are raw and personal to each one of us and the websites we’ve collected reflect that.

IIPC Collaborative Collection

When the IIPC put out a call to contribute to the Novel Coronavirus Outbreak Collaborative Collection, we got involved. Initially New Zealand sources were commenting on what was happening internationally, so URLs identified were mainly news stories, until our first reported case of Coronavirus occurred in February and then we started to see New Zealand websites created in response to Covid-19 here. We continued to contribute seed URLs to the IIPC collection, but our focus necessarily switched to the selective harvesting we undertake for the National Library’s collections.

Lockdown

The New Zealand government instituted a 4 level alert system on March 21 and we quickly moved to level 4 lockdown on March 24. The lockdown lasted a month, before gradually moving down to level 1 on June 8.

The rapidly changing alert levels were reflected in the constantly changing webpages online. It seemed that most websites we regularly harvest had content relating to Covid-19. Our selective web harvesting team focussed on identifying websites that had significant Covid-19 content or were created to cover Covid-19 events during our rapid response collecting phase. Even then it was difficult to capture all changes on a website as they responded to the different alert levels.

We were working from home during this time and connected to Web Curator Tool through our work computers. The harvesting was consistent, but our internet connections were not always stable, so we often got thrown out of the system! If we had technical issues with any particular website harvest, by the time we resolved it, the pages online had sometimes shifted to another alert level! We also used Web Recorder and Archive-It for some of our web harvests.

Due to the enormous amount of Covid-19 content being generated and because we are a very small team (along with the challenges of working from home), what we collected could really only be a very selective representation.

Unite against Covid-19 – Unite for the Recovery

Unite Against Covid-19 harvested 18 March 2020.

One prominent website captured during this time was the government website ‘Unite Against Covid-19’ which was the go-to place for anyone wanting to know what the current rules were. This website was updated constantly, sometimes several times a day.

When we entered alert level one the website changed to “Unite for the Recovery.” We expect to be collecting this site for some time. While we have completed our rapid response phase we will be continuing to collect Covid-19 related material as part of our regular harvesting.

Unite for the Recovery harvested 9 June 2020.

Economic Impact
Apart from official government websites, we captured websites that reflected the economic impact on our society, such as event cancellations and business closures. We documented how some businesses responded to the pandemic, by changing production lines from clothing to making face masks and from alcohol production to making hand sanitiser. New products like respirators and PPE (personal protective equipment) gear were also being produced. Tourism is a major industry in New Zealand and with border lockdowns still in place, advertising is now targeting New Zealanders. There is talk about extending this to a “Trans-Tasman” bubble to include Australia and possibly some Pacific Islands in the near future.

Social impact
As in many countries, community responses during lockdown provided both unique and shared experiences. New Zealanders were able to walk locally (with social distancing) so people put bears and other soft toys in the windows for kids (and adults) to count as they walked by. The daily televised 1pm Covid-19 updates from Prime Minister Jacinda Ardern and Director General of Health, Dr Ashley Bloomfield during lockdown was compulsive viewing and generated memorabilia such as T-shirts, bags and coasters. These were all reflected in the websites we collected. We also harvested personal blogs such as ‘lockdown diaries’.

Web archiving and beyond
During this rapid collecting phase, the web archivists focussed on collecting websites, and that’s reflected in this blog post. There was also a significant amount of content we wanted to collect from social media such as memes, digital posters and podcasts, New Zealand social commentary on Twitter and email from businesses and associations. This has required considerable effort from the Library’s Digital Collecting and Legal Deposit teams. You can find out more about this in an earlier National Library blog post by our Senior Digital Archivist Valerie Love. We are also working with our GLAM sector colleagues and donors to continue to build these collections.

Web Archiving at the National Library of Ireland

National Library of Ireland Reading Room © National Library of Ireland.

The National Library of Ireland has a long-standing tradition of collecting, preserving and making accessible the published and printed output of Ireland. The library is over 140 years old and we now also have rich digital collections concerning the political, cultural and creative life of Ireland. The NLI has been archiving the Irish web on a selective basis since 2011. We have over 17 TB of data in the selective web archive, openly available for research through our website.  A particular strength of our web archive is the coverage of Irish politics including a representation of every election and referendum since 2011. No longer in its infancy, the NLI has made some exciting developments in recent years. This year we have begun working with Internet Archive for our selective web archive and are looking forward to the new opportunities that this partnership will bring. We have also begun working closely with an academic researcher from a Higher Education institute in Ireland, who is carrying out network analysis on a portion of our selective data.

In 2007 and 2017, the NLI undertook domain crawling projects and there is now over 43TB of data archived from these crawls. The National Library of Ireland is a legal deposit library, entitling it to a copy of everything published in Ireland. However, unlike many countries in Europe, legal deposit legislation does not currently extend to online material so we cannot make these crawls available. Despite these barriers, the library remains committed to preserving the online story of Ireland in whatever way we can.

Revisions to the legislation are currently before the Irish parliament and if passed will result in the addition of e-publications, such as e-books, journals etc. The addition of websites to that list is currently being considered.

In 2017, the National Library of Ireland became members of the IIPC and we are excited to be attending our first General Assembly in Wellington. While we had anticipated talking about our newly available domain web archive portal and how this had impacted our selective crawls, we are looking forward to discussing the challenges we continue to face, including with Legal Deposit, and how we are developing the web archive as a whole. We may also hopefully be able to update on progress with the legislative framework.  We look forward to seeing you there in Wellington!

Archiving the Croatian web: has it been fourteen years already?

The National and University Library in Zagreb has been an IIPC member since 2008. The Croatian Web Archive (Hrvatski arhiv weba, HAW), established in 2004, is open access. The current projects include delivering metadata to Europeana, implementation of persistent identifier URN:NBN, migration to OpenWayback, development of a new user interface and integration with the Digital Library portal. Web Archiving Team has also been involved in introducing librarians, archivists and researchers to web archiving and to using HAW resources.


By Ingeborg Rudomino, Croatian Web Archive, National and University Library in Zagreb and Karolina Holub, Croatian Digital Library Development Centre, Croatian Institute for Librarianship, National and University Library in Zagreb

About HAW

The National and University Library in Zagreb (NUL) in collaboration with the University Computing Centre in Zagreb (Srce) established the Croatian Web Archive (Hrvatski arhiv weba, HAW) in 2004 and started to acquire, catalogue and archive online publications according to the legal deposit provisions of the Library Act from 1997. Due to the well-known characteristics of web resources, the NUL started to archive selectively and established selection criteria.

Fig. 1. Croatian Web Archive Homepage.

We use several methods to identify a web resource for cataloguing and archiving: the HAW team searches and browses the web; website owners or content providers fill out the Registration form or we receive notifications from the ISSN Centre for Croatia.

After identification, every resource is catalogued in the library system and automatically transferred into our custom-built archiving system, where the archiving process starts. Our long-standing experience in cataloging this type of resource has shown the process to be very challenging, and describing this dynamic and variable content results in daily interventions in the bibliographic records. Because of that, we created cataloguing guidelines with a variety of examples. Our goal has been to preserve the original websites (their look and feel) as much as possible. In order to achieve quality, each resource is approached individually during the archiving process. The DAMP software, developed by the University Computing Centre in Zagreb, was built especially for this purpose. The workflow of processing web resources is integrated within the organisational structure of the Library.

We are proud of the quantity and quality of web resources stored in the Croatian Web Archive, some of which are websites of institutions, associations, clubs, research projects, news media, portals, blogs, official websites of counties, cities, journals and books. Special attention is given to news media websites/portals, which are archived daily, weekly or monthly.

Access and the first full domain crawl

This selective approach ensures quality and provides full control over the management of web resources. So far, over 6,700 titles have been archived and almost all are publicly available. All content is full text searchable, and it’s possible to search by any word in the title, URL or keywords. Advanced search is available as well. Users can browse the HAW alphabetically and through subject categories, which are extracted from the UDC field in the catalogue.

Fig. 2. Screenshots of archived Croatian websites.

To secure permanent access to archived web resources, we have recently implemented persistent identifier URN:NBN and have assigned it to archived titles and all archived instances (Fig. 3).

Fig. 3. Screenshot of archived instances with URN:NBN.

Since 2013, the metadata from HAW is delivered to Europeana through HAW’s OAI-PMH interface.

To overcome the limitations of selective archiving, the first harvest of the whole .hr domain was conducted in 2011 with the Heritrix web crawler. Since then, we have been harvesting the .hr domain annually. The collected content is publicly available via HAW’s website through the OpenWayback access interface (Fig. 4). To date, we have conducted 7 .hr domain harvests.

Fig. 4. Screenshot of harvested website in OpenWayback.

Thematic crawls

In 2011, we started to periodically harvest websites related to topics and events of national importance using Heritrix and OpenWayback, as well. Nine thematic collections have been created, mainly related to themes such as presidential, parliament or local elections, accession to the EU and the flood in Croatia. Each collection consists of several metadata: title, size, number of seeds/URLs and description.

Training and outreach

Twice every year, we organize a workshop within the Centre of Continuing Education for Librarians. With the main goal to introduce the web archiving to library professionals and students, the workshop focuses on learning how to recognize online materials that should be preserved according to existing criteria for cataloguing and archiving Croatian web resources. The participants are also introduced to the workflow of selective archiving, .hr harvests, the process of selecting materials for thematic collections and different ways of browsing the archived content.

With the experience that we have gained throughout the years, sharing our knowledge and expertise on web archiving is something that we are happy to provide and give support to all those interested. To increase awareness about HAW and web archiving among librarians, archivists, and wider community, we try to make use of every opportunity to do so – such as presenting at national and international conferences, giving lectures to students, researchers, etc.

A few thoughts for the future

The Croatian Web Archive currently has more than 40 TB of content. We are currently working on a web interface that will have new functionalities and features including full-text search for the domain harvests and news sections for web archiving community and researchers. Also, the plan is to integrate HAW’s metadata into the Digital Library portal in order to have a single access point for all digital collections.

By combining all three approaches and using different software, the Library will attempt to cover, to the greatest extent possible, the contemporary part of Croatian cultural and scientific heritage.

Visit us: http://haw.nsk.hr/en