RE: @palewire@mastodon.palewi.re
I love this example of using Wayback Machine as a source for bringing a website back from the ashes in a more documentary form.
Preservation is access, in the future!
The web is a preservation medium.
What would we be without wishful thinking?
Salut, J'essaie d'apprendre le français!
See also: @ink@merveilles.town
RE: @palewire@mastodon.palewi.re
I love this example of using Wayback Machine as a source for bringing a website back from the ashes in a more documentary form.
I knew that Stanford's president was forced to resign due to documented evidence of him having falsified research on multiple occasions.
What I somehow missed till now was that a student at Stanford led the investigative reporting.
He now has a book about this story and the deep entanglement of the university with bigtech money & power. Here he is being interviewed on the PBS News Hour this evening:
For a project at work we wanted to be able to generate "Gold Rush" keys for MARC records, to help identify shared library holdings.
I found a Python implementation bundled up in some other code from Princeton University Library, and extracted the relevant bits to a new installable module.
https://gitlab.com/pymarc/goldrush
I gave them copyright credit, and contacted them to see if they are ok with it. If they don't like the idea I will reimplement it.
@acdha@code4lib.social for single page I really like Harvard LIL's Scoop, which I believe is used by their flagship perma.cc service via a celery job that talks to scoop-api, a (closed source?) web service that wraps Scoop:
https://github.com/harvard-lil/scoop#readme
For more than one page crawls I think that @webrecorder@digipres.club's browsertrix-crawler is still the best thing out there. I helped write this howto, which I stand by:
https://sciop.net/docs/scraping/webpages/
PS. thanks for asking me about my favorite $work thing :-)
@acdha@code4lib.social nice, happy to jump on a zoom sometime to chat about it -- browsertrix-crawler gets a lot of attention from Webrecorder since their https://webrecorder.net/browsertrix/ service depends on it.
Bernie is smart to have AOC by his side, because she actually gets the political & economic significance of AI.
https://www.youtube.com/live/6B2x2FrJa6w?si=LAB2iHygiVi91n4A
TIL that ORCID identifiers are available as #LinkedData :
https://gist.github.com/edsu/9be9658f9c6d300c569bae9b1016e108
I'm fiddling around with the Altmetric API (since we have access at $work). I sampled 12,109 DOIs (not random) from our database of publications by Stanford authors. 8,797 (73%) have Altmetric data. Altmetric have different categories of citations that they track: facebook, blogs, twitter, bluesky, news, etc.
@acdha@code4lib.social and yes, the llm bot defenses are a real problem for web archives. I saw in IIPC Slack yesterday that they are considering this as a topic for a technical meeting at the next conference in a few weeks (meeting may be open to remote participation).
Cloudflare has a Verified Bot registry, which I think some web archiving orgs have gotten on?
@stuartyeates@cloudisland.nz even more even more confusing! No wonder your stomach is turning :neocat_dizzy:
@stuartyeates@cloudisland.nz is that pointing at the same resource as https://viaf.org/en/viaf/158359701 ?
I've got a CD-R (actually quite a few) that theoretically contain backed up audio recordings from 2003. I don't know what software created it, and macOS doesnt identify it as ISO 9660, so it won't mount. Does anyone have any suggestions for things I could try to identify and access the data?