Data Wrangling: MSCS View from the trenches

•Télécharger en tant que PPT, PDF•

1 j'aime•579 vues

This document discusses best practices and lessons learned for data wrangling projects. It emphasizes starting projects by defining goals and intended outcomes. Common challenges discussed include inconsistent data encoding, missing fields, and unexpected data issues. The document provides tips for investigating data exports, mapping fields, and using tools like MarcEdit to clean data when problems arise. The overall message is that real-world library data is messy and may vary across systems.

Technologie Business

Data Wrangling: MSCS
View from the trenches
What we've learned
Where we failed
How we succeeded

You do what?

Liaisoning between tech services, project team,
and vendors on data manipulation and display

Skills:
− Marc and ILS data migration/manipulation
− Nitty Gritty details – hows and whys
− Knowledge sharing between partners
− Investigations and Implementations
− Project management
− Meeting management

Data driven? Start at the end!

What do you really want to know?

Do you have the data to answer that?

What are you going to do with the data

What is interesting vs. what is actionable

Test out your theories!!

Data driven? Start at the end!

Comparisons across institutions – match points
Started with an OCLC reclamation project
Records Sent Returned Unresolved Updated
OCLC #
Ursus 2,100,299 13,232 171,474
Colby 474,438 373
26,334
Bowdoin 624,164
37,848
Bates 656,926
25,101
TOTALS 3,855,827 13,605
260,757

Start at the end...if your ordering out

Think about what you want to get back, make
sure it goes out.

HOW will you deal with returned data?

Can all the partners do the same things in terms
of processing?

Lists, lists, lists!
What will you in/exclude if you are extracting:
types: gov docs, serials, media, e-resources
locations: ref, off-site, reserve, special collections
status: billed, missing, suppressed, withdrawn (!)
use: circ, internal use, reserves
What constitutes a circulating copy?
How are the above encoded?
Can you get what you want?

Circ Data

How long has it been retained?

Any tech processing that included circing?

Has it ever been cleared?

(… and what does it really tell you ...)

Know your vendor / programmer

What exactly is going to happen to the data,
and what will be in(ex)cluded?

Leader bib level m , s

Gov Doc? (008 / 28) ?

Printed material? Media?

Can you get it out?
Export Tables

What exactly is exported

What do they do with weird data? (b b, b 930)

Do the add any data? v.v.29 , oclc prefix

Formats of dates

Your data may vary
35109002285482 3510900228549

Document!!! REALLY!!!

Export tables and field mappings

Locations

List creation criteria

Record ranges exported and dates

Files

… a few of the ugly things we saw...

Multiple fields used for internal use (INTL
USE, COPY USE, and IUSE3)

Records with multiple 001s

Records with multiple barcodes, duplicate
barcodes, bound with items

Barcodes in 949 not 'b'

Records with no 260

3 0000003 ocm3 3_

Your data through different lenses
Points of departure:
-Merged 001s
-FRBR
-Volume vs Title counts
-Unique vs Holdings counts
-Date of data used
-Definition of public domain

When things go wrong
MarcEdit is your friend!

One more reason to thank Terry Reese
SELECT T0xx.field_data
FROM T0xx, T9xx
WHERE T9xx.field = '945'
AND T9xx.subfield = "f"
AND T9xx.field_data > 0
AND T0xx.cid = T9xx.cid
AND T0xx.field = '001'

Data Wrangling: MSCS Side
Closing Haiku:
Data is messy
While it can be normalized
Nothing is perfect

Contenu connexe

Similaire à Data Wrangling: MSCS View from the trenches

Data Exploration & BI

Cristian Guajardo-Garcia

Brandon Wong, Lead Software Engineer, Academy of Motion Picture Arts and Sciences Business Intelligence is a technology-driven process that analyzes data and forms conclusions to help assist workers to make informed business decisions. From collecting to cleaning, to morphing, to displaying we will address the pain points, tips, and tricks on how to navigate this process of converting data from raw material to a final product. You'll learn: From a high level, the process of bringing data from the "back" to the "front". Tools and best practices for cleaning and displaying data. Understanding the foundations of business intelligence to better execute on objectives. The various ways of displaying data depends on circumstance.

Data Con LA 2022 - Demystifying the Art of Business Intelligence and Data Ana...

Data Con LA

Elementary Data Analysis with MS excel_Day-1

Redwan Ferdous

BPM2017 - Integrated Modeling and Verification of Processes and Data Part 1: ...

Faculty of Computer Science - Free University of Bozen-Bolzano

Getting from raw data to deploying data-driven solutions requires technology, data, and people. All of which exist. So why aren’t we seeing more truly data-driven companies: what's missing and why? During Strata Hadoop World Singapore 2015, Pauline Brown, Director of Marketing at Dataiku, explains how lack of collaboration is what is keeping companies from building and deploying data products effectively. Learn more about Dataiku and Data Science Studio: www.dataiku.com

The 3 Key Barriers Keeping Companies from Deploying Data Products

Dataiku

Verification of Data-Aware Processes at ESSLLI 2017 1/6 - Introduction and Mo...

Faculty of Computer Science - Free University of Bozen-Bolzano

As a machine learning practitioner, you probably have met people asking the question: how can I use machine learning to solve my problem? In this talk, we'll present a few of the challenges of setting up a machine learning pipeline in the real world. We'll explain why it is fundamentally different from a typical software engineering pipeline. And we'll (try to) give a few best practices to help software engineers "think ML" and prepare their collaboration with data scientists. Recording: https://youtu.be/TZOWthpeqUY?si=MxQfT9FhPSx7fc1X&t=481

On Machine Learning Readiness

Anne-Marie Tousch

Many data pipelines share common characteristics and are often built in similar but bespoke ways, even within a single organisation. In this talk, we will outline the key considerations which need to be applied when building data pipelines, such as performance, idempotency, reproducibility, and tackling the small file problem. We’ll work towards describing a common Data Engineering toolkit which separates these concerns from business logic code, allowing non-Data-Engineers (e.g. Business Analysts and Data Scientists) to define data pipelines without worrying about the nitty-gritty production considerations. We’ll then introduce an implementation of such a toolkit in the form of Waimak, our open-source library for Apache Spark (https://github.com/CoxAutomotiveDataSolutions/waimak), which has massively shortened our route from prototype to production. Finally, we’ll define new approaches and best practices about what we believe is the most overlooked aspect of Data Engineering: deploying data pipelines.

Best Practices for Building and Deploying Data Pipelines in Apache Spark

Databricks

Level of Information Need + ISO 19650 Appointment Workflow (and eSignature fo...

Clive Jordan - fighter of Evil BIM

LoQutus: A deep-dive into Microsoft Power BI

LoQutus

Data preprocessing.pdf

sankirtishiravale

Data Warehouse Project Report

Tom Donoghue

Intelligent Data Extraction, Turning Content into Data, A Look at Advanced Ca...

DocuFi, offering HAI and Infection Prevention Analytics

The presentation will describe methods for discovering interesting and actionable patterns in log files for security management without specifically knowing what you are looking for. This approach is different from "classic" log analysis and it allows gaining an insight into insider attacks and other advanced intrusions, which are extremely hard to discover with other methods. Specifically, I will demonstrate how data mining can be used as a source of ideas for designing future log analysis techniques, that will help uncover the coming threats. The important part of the presentation will be the demonstration how the above methods worked in a real-life environment.

Log Mining: Beyond Log Analysis

Anton Chuvakin

Business Intelligence_IADC.docx

mary magdaline

What Hiring Managers Look For in Data Candidates

Ruben Kogel

About the webinar The Internet is a rich source of data, mainly textual data. But making use of huge quantities of data is a complex and time-consuming task. NLP can help with this problem through the use of Named Entity Recognition systems. Named entities are terms that refer to names, organizations, locations, values etc. NER annotates texts – marking where and what type of named entities occurred in it. This step significantly simplifies further use of such data, allowing for easy categorization of documents, analyze sentiments, improving automatically generated summaries etc. Further, in many industries, the vocabulary keeps changing and growing with new research, abbreviations, long and complex constructions, and makes it difficult to get accurate results or use rule-based methods. Named Entity Recognition and Classification can help to effectively extract, tag, index, and manage this fast and ever-growing knowledge. Through this webinar, we will understand how NER can be used to extract key entities from large volumes of text data What you will learn - How organizations are leveraging Named Entity Recognition across various industries - Live demo - Identify & classify complex terms & with NERC (Named Entity Recognition & Categorization) - Best practice to automate machine learning models in hours not months

How to analyze text data for AI and ML with Named Entity Recognition

Skyl.ai

basis data 02.pptx

MuhammadNaufalMuthah

Acc 340 Preview Full Course

fasthomeworkhelpdotcome

Talend Community Use Group Bristol: Preparing your business for mastering dat...

KETL Limited

Similaire à Data Wrangling: MSCS View from the trenches (20)

Data Exploration & BI

Data Con LA 2022 - Demystifying the Art of Business Intelligence and Data Ana...

Elementary Data Analysis with MS excel_Day-1

BPM2017 - Integrated Modeling and Verification of Processes and Data Part 1: ...

The 3 Key Barriers Keeping Companies from Deploying Data Products

Verification of Data-Aware Processes at ESSLLI 2017 1/6 - Introduction and Mo...

On Machine Learning Readiness

Best Practices for Building and Deploying Data Pipelines in Apache Spark

Level of Information Need + ISO 19650 Appointment Workflow (and eSignature fo...

LoQutus: A deep-dive into Microsoft Power BI

Data preprocessing.pdf

Data Warehouse Project Report

Intelligent Data Extraction, Turning Content into Data, A Look at Advanced Ca...

Log Mining: Beyond Log Analysis

Business Intelligence_IADC.docx

What Hiring Managers Look For in Data Candidates

How to analyze text data for AI and ML with Named Entity Recognition

basis data 02.pptx

Acc 340 Preview Full Course

Talend Community Use Group Bristol: Preparing your business for mastering dat...

Plus de Maine_SharedCollections

Dismantling Silos to Build Robust Shared Print Projects

Maine_SharedCollections

Building a Shared Print Network in New England and Beyond

Maine_SharedCollections

Developing a State-Wide Retention Policy

Maine_SharedCollections

Strategies for Preserving Maine's Collection in Print

Maine_SharedCollections

An Introduction to Maine Shared Collections

Maine_SharedCollections

Statewide Collection Analysis in Maine

Maine_SharedCollections

An Introduction to Maine Shared Collections

Maine_SharedCollections

MSCC Minerva Users Council Presentation

Maine_SharedCollections

Moving Shared Print to the Network Level

Maine_SharedCollections

How Can Digital Collections Support Shared Print Initiatives?

Maine_SharedCollections

To retain, or not retain, that is the question

Maine_SharedCollections

National Monograph Strategy

Maine_SharedCollections

“Selecting for Sustainability”Maine Shared Collections Strategy

Maine_SharedCollections

Data to Decisions: Shared Print Retention in Maine

Maine_SharedCollections

Slides from the October 21st, 2013 presentation given by MSCS Program Manager Matthew Revitt and Project PI Deb Rollins at the 2013 New England Library Association Annual Conference in Portland, ME. The session was jointly sponsored by The Academic Libraries Section (ALS) and the New England Technical Services Librarians (NETSL). A copy of the handout can be found here: http://www.maineinfonet.org/mscs/wp-content/uploads/MSCS-NELA-Handout.pdf

United We Stand: A Collaborative Approach to Legacy Print Collections

Maine_SharedCollections

Maine Shared Collections Strategy Print Archive Network Update

Maine_SharedCollections

Maine Shared Collections Strategy: Origins, Vision, Goals

Maine_SharedCollections

Communication, Project Management & Decision-Making

Maine_SharedCollections

Collaborating to Preserve Our Print Collections

Maine_SharedCollections

Managing the collective collection - print books in maine

Maine_SharedCollections

Plus de Maine_SharedCollections (20)

Dismantling Silos to Build Robust Shared Print Projects

Building a Shared Print Network in New England and Beyond

Developing a State-Wide Retention Policy

Strategies for Preserving Maine's Collection in Print

An Introduction to Maine Shared Collections

Statewide Collection Analysis in Maine

An Introduction to Maine Shared Collections

MSCC Minerva Users Council Presentation

Moving Shared Print to the Network Level

How Can Digital Collections Support Shared Print Initiatives?

To retain, or not retain, that is the question

National Monograph Strategy

“Selecting for Sustainability”Maine Shared Collections Strategy

Data to Decisions: Shared Print Retention in Maine

United We Stand: A Collaborative Approach to Legacy Print Collections

Maine Shared Collections Strategy Print Archive Network Update

Maine Shared Collections Strategy: Origins, Vision, Goals

Communication, Project Management & Decision-Making

Collaborating to Preserve Our Print Collections

Managing the collective collection - print books in maine

Dernier

Accelerating FinTech Innovation: Unleashing API Economy and GenAI Vasa Krishnan, Chief Technology Officer - FinResults Apidays New York 2024: The API Economy in the AI Era (April 30 & May 1, 2024) ------ Check out our conferences at https://www.apidays.global/ Do you want to sponsor or talk at one of our conferences? https://apidays.typeform.com/to/ILJeAaV8 Learn more on APIscene, the global media made by the community for the community: https://www.apiscene.io Explore the API ecosystem with the API Landscape: https://apilandscape.apiscene.io/

Apidays New York 2024 - Accelerating FinTech Innovation by Vasa Krishnan, Fin...

apidays

DBX First Quarter 2024 Investor Presentation

Dropbox

CNIC Information System with Pakdata Cf In Pakistan

danishmna97

MINDCTI Revenue Release Quarter One 2024

MIND CTI

EMPOWERMENT TECHNOLOGY GRADE 11 QUARTER 2 REVIEWER

MadyBayot

Keynote 2: APIs in 2030: The Risk of Technological Sleepwalk Paolo Malinverno, Growth Advisor - The Business of Technology Apidays New York 2024: The API Economy in the AI Era (April 30 & May 1, 2024) ------ Check out our conferences at https://www.apidays.global/ Do you want to sponsor or talk at one of our conferences? https://apidays.typeform.com/to/ILJeAaV8 Learn more on APIscene, the global media made by the community for the community: https://www.apiscene.io Explore the API ecosystem with the API Landscape: https://apilandscape.apiscene.io/

Apidays New York 2024 - APIs in 2030: The Risk of Technological Sleepwalk by ...

apidays

In this keynote, Asanka Abeysinghe, CTO,WSO2 will explore the shift towards platformless technology ecosystems and their importance in driving digital adaptability and innovation. We will discuss strategies for leveraging decentralized architectures and integrating diverse technologies, with a focus on building resilient, flexible, and future-ready IT infrastructures. We will also highlight WSO2's roadmap, emphasizing our commitment to supporting this transformative journey with our evolving product suite.

Platformless Horizons for Digital Adaptability

WSO2

FWD Group - Insurer Innovation Award 2024

The Digital Insurer

Strategies for Landing an Oracle DBA Job as a Fresher

Remote DBA Services

💉💊+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHABI}}+971581248768 +971581248768 Mtp-Kit (500MG) Prices » Dubai [(+971581248768**)] Abortion Pills For Sale In Dubai, UAE, Mifepristone and Misoprostol Tablets Available In Dubai, UAE CONTACT DR.Maya Whatsapp +971581248768 We Have Abortion Pills / Cytotec Tablets /Mifegest Kit Available in Dubai, Sharjah, Abudhabi, Ajman, Alain, Fujairah, Ras Al Khaimah, Umm Al Quwain, UAE, Buy cytotec in Dubai +971581248768''''Abortion Pills near me DUBAI | ABU DHABI|UAE. Price of Misoprostol, Cytotec” +971581248768' Dr.DEEM ''BUY ABORTION PILLS MIFEGEST KIT, MISOPROTONE, CYTOTEC PILLS IN DUBAI, ABU DHABI,UAE'' Contact me now via What's App…… abortion Pills Cytotec also available Oman Qatar Doha Saudi Arabia Bahrain Above all, Cytotec Abortion Pills are Available In Dubai / UAE, you will be very happy to do abortion in Dubai we are providing cytotec 200mg abortion pill in Dubai, UAE. Medication abortion offers an alternative to Surgical Abortion for women in the early weeks of pregnancy. We only offer abortion pills from 1 week-6 Months. We then advise you to use surgery if its beyond 6 months. Our Abu Dhabi, Ajman, Al Ain, Dubai, Fujairah, Ras Al Khaimah (RAK), Sharjah, Umm Al Quwain (UAQ) United Arab Emirates Abortion Clinic provides the safest and most advanced techniques for providing non-surgical, medical and surgical abortion methods for early through late second trimester, including the Abortion By Pill Procedure (RU 486, Mifeprex, Mifepristone, early options French Abortion Pill), Tamoxifen, Methotrexate and Cytotec (Misoprostol). The Abu Dhabi, United Arab Emirates Abortion Clinic performs Same Day Abortion Procedure using medications that are taken on the first day of the office visit and will cause the abortion to occur generally within 4 to 6 hours (as early as 30 minutes) for patients who are 3 to 12 weeks pregnant. When Mifepristone and Misoprostol are used, 50% of patients complete in 4 to 6 hours; 75% to 80% in 12 hours; and 90% in 24 hours. We use a regimen that allows for completion without the need for surgery 99% of the time. All advanced second trimester and late term pregnancies at our Tampa clinic (17 to 24 weeks or greater) can be completed within 24 hours or less 99% of the time without the need surgery. The procedure is completed with minimal to no complications. Our Women's Health Center located in Abu Dhabi, United Arab Emirates, uses the latest medications for medical abortions (RU-486, Mifeprex, Mifegyne, Mifepristone, early options French abortion pill), Methotrexate and Cytotec (Misoprostol). The safety standards of our Abu Dhabi, United Arab Emirates Abortion Doctors remain unparalleled. They consistently maintain the lowest complication rates throughout the nation. Our Physicians and staff are always available to answer questions and care for women in one of the most difficult times in their lives. The decision to have an abortion at the Abortion Cl

+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...

?#DUbAI#??##{{(☎️+971_581248768%)**%*]'#abortion pills for sale in dubai@

Webinar Recording: https://www.panagenda.com/webinars/why-teams-call-analytics-is-critical-to-your-entire-business Nothing is as frustrating and noticeable as being in an important call and being unable to see or hear the other person. Not surprising then, that issues with Teams calls are among the most common problems users call their helpdesk for. Having in depth insight into everything relevant going on at the user’s device, local network, ISP and Microsoft itself during the call is crucial for good Microsoft Teams Call quality support. To ensure a quick and adequate solution and to ensure your users get the most out of their Microsoft 365. But did you know that ‘bad calls’ are also an excellent indicator of other problems arising? Precisely because it is so noticeable!? Like the canary in the mine, bad calls can be early indicators of problems. Problems that might otherwise not have been noticed for a while but can have a big impact on productivity and satisfaction. Join this session by Christoph Adler to learn how true Microsoft Teams call quality analytics helped other organizations troubleshoot bad calls and identify and fix problems that impacted Teams calls or the use of Microsoft365 in general. See what it can do to keep your users happy and productive! In this session we will cover - Why CQD data alone is not enough to troubleshoot call problems - The importance of attributing call problems to the right call participant - What call quality analytics can do to help you quickly find, fix-, and prevent problems - Why having retrospective detailed insights matters - Real life examples of how others have used Microsoft Teams call quality monitoring to problem shoot problems with their ISP, network, device health and more.

Why Teams call analytics are critical to your entire business

panagenda

Dubai, known for its towering skyscrapers, luxurious lifestyle, and relentless pursuit of innovation, often finds itself in the global spotlight. However, amidst the glitz and glamour, the emirate faces its own set of challenges, including the occasional threat of flooding. In recent years, Dubai has experienced sporadic but significant floods, disrupting normalcy and posing unique challenges to its infrastructure. Among the critical nodes in this bustling metropolis is the Dubai International Airport, a vital hub connecting the world. This article delves into the intersection of Dubai flood events and the resilience demonstrated by the Dubai International Airport in the face of such challenges.

Rising Above_ Dubai Floods and the Fortitude of Dubai International Airport.pdf

Orbitshub

How to Troubleshoot Apps for the Modern Connected Worker

ThousandEyes

Artificial Intelligence Chap.5 : Uncertainty

Khushali Kathiriya

Sidekick Solutions uses Bonterra Impact Management (fka Social Solutions Apricot) and automation solutions to integrate data for business workflows. We believe integration and automation are essential to user experience and the promise of efficient work through technology. Automation is the critical ingredient to realizing that full vision. We develop integration products and services for Bonterra Case Management software to support the deployment of automations for a variety of use cases. This video focuses on the deployment of external web forms using Jotform for Bonterra Impact Management. This solution can be customized to your organization’s needs and deployed to support the common use cases below: - Intake and consent - Assessments - Surveys - Applications - Program registration Interested in deploying web form automations for Bonterra Impact Management? Contact us at sales@sidekicksolutionsllc.com to discuss next steps.

Web Form Automation for Bonterra Impact Management (fka Social Solutions Apri...

Jeffrey Haguewood

MS Copilot expands with MS Graph connectors

Nanddeep Nachan

Passkeys: Developing APIs to enable passwordless authentication Cody Salas, Sr Developer Advocate | Solutions Architect - Yubico Apidays New York 2024: The API Economy in the AI Era (April 30 & May 1, 2024) ------ Check out our conferences at https://www.apidays.global/ Do you want to sponsor or talk at one of our conferences? https://apidays.typeform.com/to/ILJeAaV8 Learn more on APIscene, the global media made by the community for the community: https://www.apiscene.io Explore the API ecosystem with the API Landscape: https://apilandscape.apiscene.io/

Apidays New York 2024 - Passkeys: Developing APIs to enable passwordless auth...

apidays

Scaling API-first – The story of a global engineering organization Ian Reasor, Senior Computer Scientist - Adobe Radu Cotescu, Senior Computer Scientist - Adobe Apidays New York 2024: The API Economy in the AI Era (April 30 & May 1, 2024) ------ Check out our conferences at https://www.apidays.global/ Do you want to sponsor or talk at one of our conferences? https://apidays.typeform.com/to/ILJeAaV8 Learn more on APIscene, the global media made by the community for the community: https://www.apiscene.io Explore the API ecosystem with the API Landscape: https://apilandscape.apiscene.io/

Apidays New York 2024 - Scaling API-first by Ian Reasor and Radu Cotescu, Adobe

apidays

ICT role in 21st century education and its challenges

rafiqahmad00786416

Tracing the root cause of a performance issue requires a lot of patience, experience, and focus. It’s so hard that we sometimes attempt to guess by trying out tentative fixes, but that usually results in frustration, messy code, and a considerable waste of time and money. This talk explains how to correctly zoom in on a performance bottleneck using three levels of profiling: distributed tracing, metrics, and method profiling. After we learn to read the JVM profiler output as a flame graph, we explore a series of bottlenecks typical for backend systems, like connection/thread pool starvation, invisible aspects, blocking code, hot CPU methods, lock contention, and Virtual Thread pinning, and we learn to trace them even if they occur in library code you are not familiar with. Attend this talk and prepare for the performance issues that will eventually hit any successful system. About authorWith two decades of experience, Victor is a Java Champion working as a trainer for top companies in Europe. Five thousands developers in 120 companies attended his workshops, so he gets to debate every week the challenges that various projects struggle with. In return, Victor summarizes key points from these workshops in conference talks and online meetups for the European Software Crafters, the world’s largest developer community around architecture, refactoring, and testing. Discover how Victor can help you on victorrentea.ro : company training catalog, consultancy and YouTube playlists.

Finding Java's Hidden Performance Traps @ DevoxxUK 2024

Victor Rentea

Dernier (20)

Apidays New York 2024 - Accelerating FinTech Innovation by Vasa Krishnan, Fin...

DBX First Quarter 2024 Investor Presentation

CNIC Information System with Pakdata Cf In Pakistan

MINDCTI Revenue Release Quarter One 2024

EMPOWERMENT TECHNOLOGY GRADE 11 QUARTER 2 REVIEWER

Apidays New York 2024 - APIs in 2030: The Risk of Technological Sleepwalk by ...

Platformless Horizons for Digital Adaptability

FWD Group - Insurer Innovation Award 2024

Strategies for Landing an Oracle DBA Job as a Fresher

+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...

Why Teams call analytics are critical to your entire business

Rising Above_ Dubai Floods and the Fortitude of Dubai International Airport.pdf

How to Troubleshoot Apps for the Modern Connected Worker

Artificial Intelligence Chap.5 : Uncertainty

Web Form Automation for Bonterra Impact Management (fka Social Solutions Apri...

MS Copilot expands with MS Graph connectors

Apidays New York 2024 - Passkeys: Developing APIs to enable passwordless auth...

Apidays New York 2024 - Scaling API-first by Ian Reasor and Radu Cotescu, Adobe

ICT role in 21st century education and its challenges

Finding Java's Hidden Performance Traps @ DevoxxUK 2024

Data Wrangling: MSCS View from the trenches

1. Data Wrangling: MSCS View from the trenches What we've learned Where we failed How we succeeded

2. You do what?  Liaisoning between tech services, project team, and vendors on data manipulation and display  Skills: − Marc and ILS data migration/manipulation − Nitty Gritty details – hows and whys − Knowledge sharing between partners − Investigations and Implementations − Project management − Meeting management

6. Data driven? Start at the end!  What do you really want to know?  Do you have the data to answer that?  What are you going to do with the data  What is interesting vs. what is actionable  Test out your theories!!

7. We Needed Data

8. Data driven? Start at the end!  Comparisons across institutions – match points Started with an OCLC reclamation project Records Sent Returned Unresolved Updated OCLC # Ursus 2,100,299 13,232 171,474 Colby 474,438 373 26,334 Bowdoin 624,164 37,848 Bates 656,926 25,101 TOTALS 3,855,827 13,605 260,757

9. Start at the end...if your ordering out  Think about what you want to get back, make sure it goes out.  HOW will you deal with returned data?  Can all the partners do the same things in terms of processing?

10. Lists, lists, lists! What will you in/exclude if you are extracting: types: gov docs, serials, media, e-resources locations: ref, off-site, reserve, special collections status: billed, missing, suppressed, withdrawn (!) use: circ, internal use, reserves What constitutes a circulating copy? How are the above encoded? Can you get what you want?

11. Circ Data  How long has it been retained?  Any tech processing that included circing?  Has it ever been cleared?  (… and what does it really tell you ...)

12. Know your vendor / programmer  What exactly is going to happen to the data, and what will be in(ex)cluded?  Leader bib level m , s  Gov Doc? (008 / 28) ?  Printed material? Media?

13. So, you think you know your data...

14. Can you get it out? Export Tables  What exactly is exported  What do they do with weird data? (b b, b 930)  Do the add any data? v.v.29 , oclc prefix  Formats of dates

15. Your data may vary 35109002285482 3510900228549

16. Document!!! REALLY!!!  Export tables and field mappings  Locations  List creation criteria  Record ranges exported and dates  Files

17. … a few of the ugly things we saw...  Multiple fields used for internal use (INTL USE, COPY USE, and IUSE3)  Records with multiple 001s  Records with multiple barcodes, duplicate barcodes, bound with items  Barcodes in 949 not 'b'  Records with no 260  3 0000003 ocm3 3_

18. Your data through different lenses Points of departure: -Merged 001s -FRBR -Volume vs Title counts -Unique vs Holdings counts -Date of data used -Definition of public domain

19. When things go wrong MarcEdit is your friend!

20. One more reason to thank Terry Reese SELECT T0xx.field_data FROM T0xx, T9xx WHERE T9xx.field = '945' AND T9xx.subfield = "f" AND T9xx.field_data > 0 AND T0xx.cid = T9xx.cid AND T0xx.field = '001'

21. Data Wrangling: MSCS Side Closing Haiku: Data is messy While it can be normalized Nothing is perfect

Notes de l'éditeur

Easy to say “we want detailed subject analysis and title lists” but if you don't have the staff time to review, does this really matter? Try to have a clear picture BEFORE starting the project. (Data can go stale … interest vs actionable data)
Easy to say “we want detailed subject analysis and title lists” but if you don't have the staff time to review, does this really matter? Try to have a clear picture BEFORE starting the project. (Data can go stale … interest vs actionable data)
Can you get what you want in a way that is meaningful to the vendor / programmer?
Do you have enough for it have value (some question if it has value at all..) Did it get checked out to processing? Another example is getting lists of barcodes into review file – ran into this where odd internal use data in different fields Do you really want to rely on it? – That 1980's Word Perfect manual vs. Portuguese poetry
You've decided what you want and you've pulled all your data … and ?? do you know how it's going to be processed.
Variations in cataloging practices over time and space Lots of oddities – no 260, no 001, multiple 001s …
Internal Use Circ in different field – different catalog
Sent data to three different places (again document what went where!)
Data is messy Nothing is ever perfect Please do not despair

Data Wrangling: MSCS View from the trenches

Recommandé

Recommandé

Contenu connexe

Similaire à Data Wrangling: MSCS View from the trenches

Similaire à Data Wrangling: MSCS View from the trenches (20)

Plus de Maine_SharedCollections

Plus de Maine_SharedCollections (20)

Dernier

Dernier (20)

Data Wrangling: MSCS View from the trenches

Notes de l'éditeur