Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

terru

@stuebinm@pleroma.stuebinm.eu
akkoma 3.19.0
  • Open on pleroma.stuebinm.eu
Sporadische Posts aus Zügen, denen Tische fehlen

Schiebt beruflich Variablen im Kreis
204 Followers
262 Following
11 Posts
git repos:
https://stuebinm.eu/git/
preferred pronouns:
it/its or something i dunno 🤷‍♀️
gender:
: ⊥ → a
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 5mo ago
Replying to
@jonmsterling i always suspect this is less about it actually being 'dead', and people being unaware that one can, indeed, step slightly outside of what is the current mainstream, and find useful tools with widespread support there, too. RSS is not dead, but its presence as a default-tech-giant-preinstalled-everywhere-app category is, and to some people that appears to be the same (?)
1
9
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 5mo ago
Replying to
@jonmsterling > as if that is a requirement ?? there does seem to be some core aspect of tech-bro-ism that's something like, "everything is a startup, even when it isn't", and it gets relentlessly applied to open source projects. it's very strange, and exhausting at times …
1
6
1
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 5mo ago
Replying to
@jonmsterling I assume it's some kind of SEO thing or such? But I dunno, I ignore such things when they come up 🤷
1
3
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 5mo ago
Replying to
@jonmsterling for developers, bragging rights (+ maybe better chances at getting a job, in some places which share the same mentality?) for company-owned projects, metrics to make some executive happy and show how deeply involved they are in the 'open-source community'
1
0
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 6mo ago
Replying to
@leah@blahaj.social the LLMs have Clauded their minds, one might say 🙃
1
0
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 6mo ago
Replying to
@timezone can recommend simply having all of the EU (+ a little more) downloaded; it doesn’t even take up that much space
1
1
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 6mo ago
Replying to
@firefly if you convert it into imperial units, you can finally measure how many cups you have, in cups!
1
0
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 7mo ago
event called it trans, of course there is a shark!
0
0
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 9mo ago
POV: looking at available wifi ssids on a train to #39c3
pleroma.stuebinm.eu
0
1
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 6mo ago
Ah, die neuen Garnituren für das Oktoberfest!
0
0
0
0
Open post
terru @stuebinm@pleroma.stuebinm.eu
· 5mo ago
okay, so, I looked at the state of OCR again (for "i have to many crusty-looking pdfs and maybe i can make them nicer to read" reasons)

… and now i am kinda mad about AI

because, well, my idea was, if i have a bad blurry scan of a paper but reasonable OCR of it, surely i can just replace the image by text characters which the OCR found? And it turns out, this works surprisingly well! … for prose. As soon as there's even one equation, subscript, or anything else non-trivial (or, god forbid, a table, or a *rotated* table), tesseract kinda gives up. No! Worse, it does not give up, it produces nonsense, so now there's just, symbol salad that you can't even easily detect as likely error! yay!

but!

this looks, I thought, like the kind of problem where the AI hype might've actually produced reasonable advances as a side effect (the way RNNoise is also very neat). And wouldn't you know it, they did! Throwing unlimited funding (and some machine learning) *can* produce results, it turns out!

Except, of course, they are AI-hype-brained results. I can throw crusty theoretical compsci papers at their demonstration website (I assume these were all scraped already, so I have little qualms about it — tho ig there's some risk that their were in the training set and someone had to manually do a nice version, so take the following with some salt) and it deals wonderfully with all sorts of complicated formatting stuff

… but nobody in this space is interested in using any of that to make the world a place with more accessible, searchable pdfs. They're purely interested in "linearising" the pdf into a markdown document, so it can be thrown into a machine learning pipeline. You might say, so what? I just wanted to read papers, and while I like beautiful typesetting and layout, surely I can just read it as markdown instead if i'm interested in the content? But well, they also discard anything that does not fit into markdown's text model. Bullet lists, tables? Nicely converted. Diagrams? Images? Figures? Simply gone. And the model does not tell one where the text came from, just the text itself, so if you just look at the generated markdown you might not even know you've missed the most important part

and just … i dunno, i'm frustrated. there might be something close in this conceptual space that would genuinely help make tools better. But instead we get things that willfully discard information and whose main metric is "dollars per million pages read", to be put in a paper that reads like a marketing press release

:BlobhajSadReach:
0
3
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 04:21:16 UTC