Data Understanding in CPMAI

Data Understanding in CPMAI
Table of Contents

Phase 2: Data Understanding in CPMAI

Data Understanding is the second step in the CPMAI methodology. During this step, teams must gather, shape, and analyze relevant data needed for the AI solution. Teams must also ensure that the organization’s data can adequately support the project defined in Business Understanding. This phase concludes with the submission of a data assessment report. Data Understanding is a foundational phase. Building AI solutions without this phase typically creates significant design inefficiencies, as it is often discovered during this phase that the data needed does not exist or that, if it does, it exists in a form that is unusable.

What Is the Purpose of Data Understanding in CPMAI?

Data Understanding is an opportunity to work the Business Case (defined during Phase 1: Business Understanding) within the reported Data realities. An AI project can be well-scoped, however, it is still a failed project if the data supporting the project does not exist or is inaccessible or if it does exist, then it is in a form that is unusable. This phase serves as an opportunity to rethink or pause the project if it becomes clear during this analysis that it is necessary to avoid proceeding to the costlier phases of Data Preparation and Model Development.

What Activities Occur during Data Understanding?

Activities that occur during Data Understanding include: understanding data requirements, mapping data sources, identifying data subject matter experts, building partner alignment, and drafting data assessment reports. Each activity closes the gap between the data a project is designed to use, and the data a project organization is able to provide.

Activities

Products

Data requirements documentation

Documentation of the project data types, formats, and ballpark data volumes

Data source mapping

A record captures the location of the required data within the organization

SME identification

A list of records captures persons within the organization who can articulate the meaning and context of the data

Confirmed infrastructure

Verification of available workspace and storage infrastructure for the data processing pipeline

Data assessment report

An assessment of available data capturing both identified quality and quantity issues

How Are Data Requirements Documented?

Data requirements document datasets (i.e. types, characteristics, volume) that the AI solution will use to address a given business opportunity or challenge. The transformation of business problems and success criteria to unambiguous and technically compliant requirements occurs in Phase 1. For example: “last 24 months of customer purchase data” is not a sufficient requirement; the requirement should include, “last 24 months of customer purchase data at the transaction level including timestamp, SKU, customer ID.”
Each of the aspects of Business Understanding must be documented as a data requirement.
Estimate (order of magnitude) the volume of data needed to develop a satisfactory model.

.

What Is Involved in Mapping Data Sources?

Mapping data sources is the identification of where data, as defined above, logically or physically is known to exist (the internal or external company databases, third-party data vendors, publicly available data, data that is manually captured, etc.), and confirming that the data is available for harvesting. Teams assume data “is somewhere in the organization”, but they do not research the system housing the data, and if access to the data formally is required, a data sharing contract must be executed.
Note every company internal system, database, or application that may have relevant data.
Note publicly available data that can be used to supplement internal data.
Check each data source for the documentation of ownership and the process that the user goes through to access the data.
Circle any data source requiring a formal agreement, procurement, or compliance review prior to use.

How Do You Find Data Subject Matter Experts?

It is important to understand what data means to a person, as data on its own is relatively meaningless. For instance, while legacy systems exist, the same field labeled “status” may have different meanings to dozens of different employees.
Identify the data source owner for each data source noted while mapping.
Confirm each data source owner precisely describes how data was collected, what quality issues affect data, and how the schema has changed.
Meet with data source owners and plan for the Data Preparation stage prior to starting Data Cleaning to avoid quality issues.
Informal knowledge of data can supplement an empty data dictionary, so document informal knowledge while it is available.
Coordinating Infrastructure Requirements for Your AI Workspace
Infrastructure means the company provides the data storage, the compute resources, and the project workspace prior to the start of data preparation. Not having the required resources in the middle of a project causes the most delays and problems in building the required AI systems. This refers to the required data infrastructure (lakes, compute, ML) and the data scientists’ access rights.
Verify that the storage space is sufficient and that the data format is supported.
Verify that enough compute resources are available and that budget is allocated for data preparation and model training.
Do not provide access to the technical infrastructure after the permissions have been scheduled; provide access before permissions are scheduled.
Find out if there are going to be security or compliance reviews that will allow sensitive data to be used in the working environment.

What do we look for in quality data and compliance?

When assessing quality and compliance for data, we need to see if data has sufficient rigor and privacy control, as there are reliability and legal issues that occur with the use of models trained on low quality data or non-compliant data after the model is deployed. This step captures the concepts behind CPMAI’s responsible-AI principles, which cover the entire lifecycle of the data, and most importantly, starts here with data and prior to the development of the model.

What happens if Data Understanding is skipped or rushed?

Data Understanding is a process that evaluates the quality and quantity of data to determine if data meets acceptable standards for the project. Data Understanding is a critical step in the process to advance, improve, or halt a project and is documented in a data assessment report. Data Understanding is a process that assesses available data, quality, quantity, and any issues for Data Preparation.

What happens if Data Understanding is skipped or rushed?

When we skip or rush through Data Understanding, there are assumptions about the availability and quality of data, and a project starts Data Preparation and Model Development without properly understanding the data. Then, issues begin to show during data cleaning which results in delays and having to do a significant amount of work again. This is one of the more expensive gaps that exists in the CPMAI process, as a lot of teams have spent engineering time prior to even realizing the gap.

Data Preparation is started with data that is missing and results in a return to sourcing data

When data has undocumented quality issues, models are developed which result in unreliable and undesirable outcomes which Surface during Model Evaluation.

Compliance gaps are only identified when the legal or privacy reviews have completed a significant amount of work on the project.

How Does Data Understanding Relate to PMI-CPMAI Exam?

In the PMI-CPMAI exam, Data Understanding is aligned with Domain III, which covers Identify Data Needs, and about 26% of the exam. So, it is just as important as Domain II. Data Understanding is closely associated with scenario-based questions that deal with mapping data sources, reviewing data readiness and ensuring compliance. These are the major components of the domain as described in the PMI Exam Content Outline. .

Where Does Data Understanding Fit in the Full CPMAI Methodology?

Data Understanding Is the second of the six phases in the CPMAI methodology. For additional information and understanding of other phases of the methodology in relation to Data Understanding, see The CPMAI Methodology: 6 Phases Explained.

Where Can You Develop Data Understanding in Practice?

PMTI provides practice-based data readiness assessment workshops as part of live, instructor-led training for the CPMAI Certification Training program.

FAQs Phase 2: Data Understanding  

What do we call the output of Data Understanding?

A report on the status of data (if it exists) and on the expected quality of the data and gaps, with a recommendation on how the project should proceed: by continuing the project, by requesting additional data, or by waiting to undertake the project.

What is the most significant risk from skipping Data Understanding?

In the process of Data Preparation it is possible that unrecoverable compliance gaps will be found in the data, and that the desired data does not exist, or that access to it has not been secured. This will certainly represent a significant investment of both time and money.

What other stakeholders, apart from the project manager, are involved in Data Understanding?

Data experts who understand what certain datasets represent, and also the data source owners and those with access and compliance.

Which PMI-CPMAI exam domain is Data Understanding most strongly associated with?

Domain III – Identify Data Needs, which accounts for approximately 26% of the exam.

.

Yad Senapathy
Yad Senapathy

Your project managers will be trained on the PMI PMBOK Guide's best practices and ethics. They'll understand the framework of a successful project from initiating to close.

Share this article
Twitter
Facebook
Linkedln
whatsapp
telegram
pinterest
Get in Touch With Us