Skip to content
This repository has been archived by the owner on Oct 27, 2022. It is now read-only.

Latest commit

 

History

History
23 lines (22 loc) · 2.2 KB

2022-07-07-ana-trisovic.md

File metadata and controls

23 lines (22 loc) · 2.2 KB
layout title author rse excerpt date external_media length duration explicit resources
post
Joys and Challenges with Big Research Data
@vsoch
Ana Trisovic
Data engineering is an increasingly important part of research software engineering, and Ana tells us why.
2022-07-07 8:30:00
20193947
00:34:01
no
name url
Ana's Website
name url
On Twitter
name url
The Dataverse Project
name url
National Studies on Air Pollution and Health

Ana Trisovic is a Research Associate at Harvard School of Public Health and a Sloan Fellow at the Institute for Quantitative Social Science. Effectively, she does data engineering for her research group and works on reproducible data and software dissemination. First, Ana speaks of her background, from her first job at Microsoft Development Center Serbia, to CERN, UChicago, and Harvard. She shares what inspired her to pursue projects relating to open-source software, open data, and open science. Her work focuses on big data workflows and research reproducibility, and she shares her experiences working with particle physics experimental data, geospatial and climate data, and sensitive medical data. As a member of Consortium of Scientific Software Registries and Repositories (SciCodes), she contributes to research data and software sharing and preservation efforts. Her study shows that research software and code scripts are frequently shared with data, and she is working on better supporting those in the Dataverse data repository. We discuss data engineering roles in the broader RSE scope and recognize them as undervalued yet critical for research groups working on secondary data analysis. Ana speaks of the joys and challenges of working with diverse datasets and the value of open-source software, reusable data workflows, and adequate documentation. She shares recommendations for publishing research data with software and emphasizes the role of data repositories. We end the conversion with community engagement topic ideas.