Elektrine
Log in Register
Paige Chat Timeline Gallery Friends Email Drive DNS Private DNS Domains VPN Kairo Nerve
Remote

Daniel van Strien

@danielvanstrien@sigmoid.social
mastodon 4.7.3
  • Open on sigmoid.social

📖🤗 Machine learning Librarian at Hugging Face

309 Followers
354 Following
24 Posts
Joined December 19, 2022
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 24mo ago

ColPali is revolutionizing multimodal retrieval. Can we make it even more effective with domain-specific fine-tuning?

Check out my latest blog post, where I create a dataset for fine-tuning a ColPali model for a new domain using an open Vision Language Model.

https://danielvanstrien.xyz/posts/post-with-code/colpali/2024-09-23-generate_colpali_dataset.html

danielvanstrien.xyz
2
0
3
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 21mo ago

Was 2024 the year of datasets? Is 2025 the year for community-built datasets?

It's exciting to see the progress of many languages in FineWeb-C:
- Total annotations submitted: 41,577
- Languages with annotations: 106
- Total contributors: 363

1
0
1
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 25mo ago

Can we search for datasets on the @huggingface Hub based on their content?

> Some datasets lack good documentation 😢
> The dataset viewer preview offers a wealth of information

🤔 How about: query -> dataset based on structure content?

Check out V1: https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search

huggingface.co
1
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 25mo ago

The @huggingface's Semantic Dataset Search is back in action! Find similar datasets by ID or do a semantic search of dataset cards.

Give it a try:
https://huggingface.co/spaces/librarian-bots/huggingface-datasets-semantic-search

huggingface.co
1
0
2
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 29mo ago

Created an "Awesome Synthetic Datasets" list in my ongoing quest to learn more about building synthetic datasets using large language models. Currently includes important tools, datasets, and papers.

Check it out here: https://github.com/davanstrien/awesome-synthetic-datasets

github.com
1
0
1
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 29mo ago

Translations from 56 contributors, based on a dataset by 314 community members! These translations will facilitate the creation of evaluations, experimentation with SPIN, building DPO datasets, and more. Interested in contributing to datasets? https://github.com/huggingface/data-is-better-together

github.com
1
0
1
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 30mo ago

As part of the Multilingual Prompt Evaluation Project (MPEP), we are now automatically exporting the @argilla_io datasets to the @huggingface Hub. We have more than 15 active community-led translation efforts collaborating to enhance datasets for various languages. ❤️

1
2
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 30mo ago

Doing some rare "front end" work 😬

1
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 32mo ago
Replying to
@mia@hcommons.social Just added a note/link to BL :)
1
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 25mo ago
Replying to
You can help improve this project by rating synthetic user search queries for hub datasets. If you have a @huggingface login, you can start annotating in @argilla_io in < 5 seconds here: https://davanstrien-my-argilla.hf.space/dataset/1100a091-7f3f-4a6e-ad51-4e859abab58f/annotation-mode
davanstrien-my-argilla.hf.space
0
1
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 32mo ago
Replying to
While the availability of public domain newspapers in the UK may be smaller than in other European countries, this dataset still contains many tokens. It can be utilized to diversify training data and conduct research on historical language models.
0
2
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 25mo ago
Replying to
I need to do some tidying, but I'll share all the code and in-progress datasets for this soon!
0
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 32mo ago
Replying to
You can also find more historic English tokens in the blbooks dataset, a dataset of public domain digitised books, which consists of around 7.67 billion tokens. https://huggingface.co/datasets/biglam/blbooks-parquet
huggingface.co
0
1
1
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 23mo ago

Researchers: Want your ML datasets to have more impact? Share them on @huggingface Hub!

✨ Benefits:
• Visibility in the ML community
• Interactive data viewer
• Support for TB-scale datasets
• Integration with @DataPolars @pandas_dev @duckdb and more
https://huggingface.co/blog/researcher-dataset-sharing

huggingface.co
0
0
2
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 26mo ago

Is your summer reading list still empty? Curious if an LLM can generate a book blurb you'd enjoy and help build a KTO preference dataset at the same time?

A demo using @huggingface Spaces and @gradio to collect LLM output preferences: https://huggingface.co/spaces/davanstrien/would-you-read-it

huggingface.co
0
3
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 27mo ago

SPIQA from @Google is a large-scale question-answering dataset centred on figures, tables, and text paragraphs from scientific research papers in various computer science domains.
https://huggingface.co/datasets/google/spiqa

huggingface.co
0
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 28mo ago

HelpSteer2 from @nvidia is an open-source dataset to train top-performing reward models!
- 21,362 samples with annotated attributes
- Attributes: Helpfulness, Correctness, Coherence, Complexity, Verbosity
- Multi-turn prompts
- 88.8% on RewardBench
https://huggingface.co/datasets/nvidia/HelpSteer2

huggingface.co
0
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 29mo ago

Who wants to be 100!?
https://github.com/huggingface/data-is-better-together

github.com
0
2
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 30mo ago

VISION2UI: A Real-World Dataset with Layout for Code Generation from UI Designs: https://huggingface.co/datasets/xcodemind/vision2ui

huggingface.co
0
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 30mo ago

Experimenting with TL;DR summaries for @huggingface datasets using a Chrome plugin.

0
0
0
0
Open post
Daniel van Strien @danielvanstrien@sigmoid.social
· 26mo ago
Replying to
@arnicas Occasionally the books sounds interesting but often the blurbs are not very good. Think LLMs are still very lacking in this kind of task tbh.
0
1
0
0
Back
313k7r1n3
Elektrine

Tor hidden service

elekhj7afj4qnrr4yd3bkzslsyo5jgfxw3orgjkhlcxifueodybyiiad.onion

I2P eepsite

j6b6cyk6gjmepjih7jjadxgxvvf3lzzujljuu2v4biemzpg3naya.b32.i2p

Platform

  • Email
  • Chat
  • Timeline
  • VPN
  • DNS

Company

  • About
  • Contact
  • FAQ
  • Lite (no JS)

Legal

  • Terms of Service
  • Privacy Policy
  • Transparency Report
  • Report Abuse
  • Warrant Canary
  • VPN Policy

Support

  • support@elektrine.com
  • Report Security Issue
Mail client setup IMAP mail.elektrine.com:993 POP3 mail.elektrine.com:995 SMTP mail.elektrine.com:465
© 2026 Elektrine. All rights reserved. Server: 15:22:17 UTC