I think animals that hibernate for much of the year are pretty cool because what do you mean you get to only experience your favorite season?
Eric R. Scott
mastodon 4.7.3Scientific Programmer & Educator | Data Engineer with grantwitness.org | Mentor for @Posit Academy | Ecologist | Tea geek
Posts are mine and do not represent my employers in any way
If you can't help but do a lil dance when you eat something tasty, I think we'll be friends.
Migrated from fosstodon.org so I can see all my friends again!
I cannot for the life of me figure out a good way to read in a dataset made of hundreds of CSVs with *mostly* the same columns (in
). I wrote a blog post that is basically an extended reprex. Comments and solutions are welcome!
@defuneste@fosstodon.org @cd_newton@hachyderm.io The duckplyr version does though!! It works!!!
I went to my local middle eastern market today, not remembering that it is Eid. A little like Trader Joes the day before Thanksgiving except everyone was smiling and in a good mood. Anyway, Eid Mubarak to all who celebrate
*Finally* got around to testing new {purrr} and {mirai} features on the #UniversityOfArizona #HPC and it is such a big step forward for researchers who would prefer not to leave their R comfort zone like myself. Here's a gist that you can run from an interactive session to launch SLURM jobs as workers and use them for parallel computing with `purrr::map()`: https://gist.github.com/Aariq/d52a0d861e5a2c90540548caa47021af
My wife is amazing at guessing movies my mom wants to see from extremely vague descriptions.
1. The one about a woman that won an award
2. The funny space one with the tree
She got both of these right, can you?
What is the criterion that the Criterion Collection uses? Good movie?
@cd_newton@hachyderm.io @defuneste@fosstodon.org duckdb::duckdb_read_csv() (the R fucnction) does not have the union_by_name argument 😭
@nxskok@cupoftea.social Yeah, this works (I have an example of this in the blog post), but it's not ideal because the memory overhead is MUCH greater than read_csv(fnames).
@geospacedman@mastodon.social For this particular project, I think I need a database-like approach. I had high hopes for rearranging CSVs in hive style partitioning and using arrow::open_dataset(). Looks like DuckDB is going to be the way to go.
But in general, it would be nice for read_csv(files, col_select = 1:5) to "just work" even when some of the files have 6 columns, because it's a lot more performant than map(files, read_csv) |> list_rbind() (or the data.table equivalent, I'm guessing)
