About

Data Maturity at Tarkam

At Tarkam, data maturity is a structured look at how an organisation plans, collects, analyses, stores, shares, and uses data. The Data Maturity Assessment is built on the Data Maturity Matrix (DMM): a full picture of data-related processes across the organisation. The DMM follows the data pipeline — the life-cycle of data inside an organisation. That pipeline covers six stages: planning and collection; validation and cleaning; analytics; storage and archival; output artefacts; and data sharing. Each stage is reviewed for capacity, process, and challenges. You answer a set of questions about your organisation's data work. The engine maps those answers onto the matrix, assesses how mature each part of the pipeline is, and turns that into a diagnostic: recommended actions, priorities you can choose, and a roadmap for strengthening data processes.

Read the Data Maturity Assessment Engine design document.

Model

The Analytical Model

The Data Maturity Matrix is a framework to evaluate the maturity of an institution’s data practices, based on the six core elements of the data pipeline: data planning and collection; data validation and cleaning; data analytics; data storage and archival; output artefacts; and data sharing. This framework gives us a consistent way to understand an organisation’s data processes. Organisations themselves are very different — in operational requirements, risk profiles, research focus, and institutional incentives.

A distinct data maturity model is needed for organisations that work on the ground to design and implement interventions. They operate under different conditions:

  • Data is often collected in low-resource field environments
  • Staff are usually programme-related rather than technical
  • Data must support programme design, accountability, and advocacy
  • Populations may be vulnerable, raising ethical and privacy concerns
  • Infrastructure constraints such as connectivity, devices, and staff turnover

Maturity indicators must reflect these different realities. So we use different data maturity frameworks for: (a) research and large-scale implementation organisations, and (b) NGOs and CSOs.

As we work with a broader set of organisations, data maturity models for other kinds of institutions may differ slightly from the ones we present here. Even when the specifics vary, key aspects stay the same across all models: a common architecture of the data pipeline; a common logic underlying maturity (the evaluation parameters); and common maturity categories (Ad Hoc, Emerging, and Stable).

Data Maturity Model for NGOs/CSOs

The data maturity of NGOs/CSOs can be mapped on the following table:

LevelCategoryDescription
1Ad HocData collection happens, but it is project-driven and informal. It is often driven by donor reporting, staff initiative, and short-term projects. Institutional data systems do not exist.
2EmergingProgrammes collect data more consistently. Some internal systems exist but are not institutionalised.
3StableData becomes an organisational asset, not just a project requirement. Standard systems and policies exist.

A detailed breakdown of the three levels is given below.

Level 1: Ad Hoc

Core Overall Characteristics
Data collection happens but is unstructured and project-dependent. Systems depend heavily on individual staff practices. Data is mainly collected for donor reporting or documentation. There are no formal policies, governance structures, or institutional standards. Data is viewed primarily as administrative reporting material, not a strategic organisational asset.
Planning & Collection
Data is collected through surveys, field notes, and registers. Tools are created by individual staff. There are no standard questionnaires. Ethical consent is inconsistent. Common tools: paper forms, notebooks, WhatsApp, Excel.
Validation & Cleaning
Data entry errors are common. There is no cleaning protocol. Duplicate or incomplete records.
Analytics
Basic counts and summaries. Often used only for donor reports.
Storage & Retrieval
Files are stored on individual laptops or drives. Paper registers sit in field offices.
Sharing
Data is shared mainly through reports. Raw datasets are rarely shared.
Outputs
Outputs are mainly for donors and internal reporting. Formats: narrative reports, activity summaries.
Institutional Risk Profile
Very high risk of institutional memory loss. High risk of data loss, programme decision quality, and privacy breach. Very low data reliability and scalability.

Level 2: Emerging

Core Overall Characteristics
Programmes collect data more consistently. Some shared tools exist. Teams begin using digital data tools. Data is increasingly used for monitoring programmes. Systems remain project-driven rather than institutionalised.
Planning & Collection
Standard survey tools are used across projects. Mobile data collection is sometimes introduced (for example, Kobo, ODK). Staff are trained in basic data collection.
Validation & Cleaning
Data is reviewed at project level. Basic consistency checks.
Analytics
Programme monitoring indicators are used. Analysis is mostly Excel-based. Simple trend analysis.
Storage & Retrieval
Cloud storage is introduced (Google Drive and similar). Programme-level data folders. Still fragmented across projects.
Sharing
Reports and presentations are shared with donors, government partners, and networks. Some datasets are shared with partners.
Outputs
Outputs include programme reports, evaluation summaries, presentations, and case studies. The audience expands to government agencies and partner NGOs.
Institutional Risk Profile
High risk of institutional memory loss. Moderate risk of data loss, programme decision quality, and privacy breach. Moderate data reliability. Low scalability.

Level 3: Stable

Core Overall Characteristics
The organisation adopts data policies and standard workflows. Data becomes an institutional resource. Governance frameworks are introduced. Data is used actively for programme improvement and accountability.
Planning & Collection
Organisation-wide data collection protocols. Standard monitoring frameworks. Ethical consent procedures are defined. Data management plans for projects. Tools include mobile survey platforms and structured monitoring frameworks.
Validation & Cleaning
Cleaning protocols are documented. Data verification steps are built into the workflow. Field monitoring for quality checks.
Analytics
Programme analytics are used regularly. Data informs programme design decisions. Example uses: identifying vulnerable groups, improving programme targeting, measuring intervention impact.
Storage & Retrieval
A centralised data repository. Programme data is indexed and archived. Metadata standards are adopted. Clear access controls.
Sharing
Data sharing guidelines are established. Anonymisation protocols are defined. Datasets are sometimes shared with researchers, government, and collaborative platforms.
Outputs
Outputs are diversified: dashboards, infographics, community reports, policy briefs. Audiences include communities served, policymakers, and sector networks.
Institutional Risk Profile
Low risk of institutional memory loss. Low risk of data loss, programme decision quality, and privacy breach. High data reliability. Moderate to high scalability.

Data Maturity Model for research and large-scale implementation organisations

The data maturity of research organisations can be mapped on the following table:

LevelCategoryDescription
1Ad HocIndividual-driven, undocumented, reactive processes
2EmergingRepeatable but fragmented practices, project-based systems
3StableInstitutional standards, documented workflows, central repository

A detailed breakdown of the three levels is given below.

Level 1: Ad Hoc (Individual Driven Research Practice)

Core Overall Characteristics
Data processes vary entirely by project or researcher. Documentation is inconsistent or absent. Storage is fragmented across personal drives. No formal data governance. No defined lifecycle model.
Planning & Collection
Research design is strong but undocumented. No standardised Data Management Plan. Ethics are handled informally. Secondary data is used without a validation protocol.
Validation & Cleaning
Manual cleaning (Excel dominant). No cleaning logs. No audit trail. Version confusion is common.
Analytics
Spreadsheet based. Transformations are undocumented. No reproducibility. Results are embedded directly into reports.
Storage & Retrieval
Shared drives and external hard disks. No metadata. Retrieval depends on individuals.
Sharing
Report-based sharing only. Raw datasets are rarely shared. No anonymisation standard.
Outputs
Technical reports. Minimal segmentation by audience.
Institutional Risk Profile
High institutional memory loss risk. High reproducibility risk. Moderate to high privacy risk. Low scalability.

Level 2: Emerging (Repeatable but Unstandardized)

Core Overall Characteristics
Teams have consistent internal practices. Some shared tools and templates. Cloud storage is used. Awareness of governance gaps.
Planning & Collection
Conceptual rigor is strong. Tools are archived inconsistently. Secondary data reliability concerns are recognised. Ethical considerations are present but informal.
Validation & Cleaning
Repeatable cleaning practices. Some internal cross-checking. Still no institutional cleaning log requirement.
Analytics
Structured Excel workflows. Some use of qualitative software. Limited scripting. No formal analytic pipeline documentation.
Storage & Retrieval
Cloud folders organised by project. Still a fragmented architecture. No dataset registry. Version control is weak.
Sharing
Some open data sharing. No institutional licensing framework.
Outputs
Strong policy briefs and reports. Workshops and dissemination events.
Institutional Risk Profile
Moderate reproducibility risk. Moderate governance risk. Growing scale tension. Staff-dependent knowledge retention.

Level 3: Stable (Institutionalized and Standardized)

Core Overall Characteristics
A data governance framework exists. Standard templates are required. Roles are defined across the data lifecycle. Documentation is mandatory.
Planning & Collection
An institutional data management plan is required for all projects. An ethical review mechanism is established. A secondary data validation checklist. An instrument repository is maintained.
Validation & Cleaning
Documented cleaning SOP. Mandatory cleaning logs. Version control standards. A quality control checklist is applied.
Analytics
Separation of cleaning and analysis. Reproducible workflow (Excel + scripted hybrid). Visualisation standards are defined. Documented transformation records.
Storage & Retrieval
A centralised data repository. A metadata schema is required. The dataset registry is searchable. A defined access control structure.
Sharing
A data sharing policy is adopted. An anonymisation protocol is documented. A licensing framework is in place. A standardised dataset release package.
Outputs
A tiered dissemination model: Technical, Policy, Public. An audience segmentation framework.
Institutional Risk Profile
Reduced institutional memory risk. Lower privacy exposure. Moderate scalability capacity. Stronger compliance alignment.
Read the Data Maturity Model document.

Technology

Third-Party AI Processing

Loading model details…

Privacy

Tool Data Collection & Privacy

We collect operational data through this tool strictly to generate customised reports evaluating the data maturity of participating organisations. All data collected directly through the Data Maturity Assessment tool is used solely for this report generation purpose. Aggregated, non-identifiable summaries may be analysed to understand the overall state of data maturity among CSOs and to help strengthen sector-wide data practices. We strictly commit that we will not share any identifiable or sensitive data with any third party without prior explicit permission.