The CTDL export, and what it does and does not carry
This project publishes California's training programs as CTDL, the vocabulary Credential Engine maintains for describing credentials and learning opportunities. This page is the export's own account of itself: which classes and properties it fills in, what the source record says that it drops, and what a separate validator found when it was pointed at the result.
A mapping is only worth anything if somebody can check it. So the counts here are produced by the export at the moment it runs, the omissions are counted the same way as the coverage, and the validator's findings are published whichever way they came out.
What this is not
- None of this has been published to the Credential Registry. Not a submission, not a sandbox, nothing. The records exist as files in this project and nowhere else.
- This is not affiliated with, endorsed by, or reviewed by Credential Engine. They publish CTDL openly and this project reads it; that is the whole of the relationship.
- The identifiers are derived locally and are not Registry-assigned. A real CTID is issued when a resource is published to the Registry, which these are not, and the identifier URIs deliberately live on this project's own host rather than on a registry domain.
- It is a demonstration of mapping discipline rather than a production publication. It exists to show how a public source maps onto a public vocabulary, and what is lost on the way.
Note: this export describes the 2026-08-07 snapshot, and the rest of this site is currently serving 2026-09-18. The figures below describe the export, not the pages around it.
On this page
What the export contains
One entity per training program, one per distinct provider name as filed, and one statistics profile for each program that reports at least one outcome. A program that reported nothing gets no statistics profile at all: an empty one would read as "measured, and empty".
| CTDL class | Entities |
|---|---|
ceterms:LearningProgram | 3,266 |
ceterms:CredentialOrganization | 584 |
qdata:DataSetProfile | 2,057 |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
Which properties are filled in
Every property below is emitted only where the source asserted something. A blank is a blank: no placeholder, no zero, and nothing inferred from a neighboring field. The order is the export's own, not best-first.
| Property | Programs carrying it | Share |
|---|---|---|
ceterms:name | 3,266 | 100% |
ceterms:description | 3,266 | 100% |
ceterms:subjectWebpage | 1,654 | 51% |
ceterms:offeredBy | 3,266 | 100% |
ceterms:occupationType | 3,266 | 100% |
ceterms:estimatedCost | 3,265 | 100% |
qdata:relevantDataSet | 2,057 | 63% |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
Which outcome measures are projected
Each reported measure becomes one metric and one observation inside the program's statistics profile. The share is against the programs that have a statistics profile at all, not against every program: a measure missing because a program reported nothing is a different fact from a measure missing from a program that reported something else.
| Measure | Observations | Of programs with any outcome |
|---|---|---|
| Median earnings in the second quarter after leaving | 1,384 | 67% |
| Credentials earned | 1,800 | 88% |
| Employed two quarters after leaving (count) | 1,766 | 86% |
| Completion rate | 2,047 | 100% |
| Employment rate two quarters after leaving | 1,760 | 86% |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
What the export does not carry
The source record says more than this export projects. Counting only what was emitted would describe a projection as though it were the whole record, so the dropped fields are counted too, with the CTDL term that would have carried each one where such a term exists.
Where a CTDL term is named, the vocabulary has somewhere to put this and the export does not use it. That is a gap in the export, not a limit of CTDL, and it is stated that way rather than left for a reader to work out from an absence.
| What the source says | Programs reporting it | CTDL term that would carry it |
|---|---|---|
| The CIP code for the field of study | 3,266 | ceterms:instructionalProgramType |
| Online, in person, or both | 3,266 | ceterms:learningDeliveryType |
| How long the program takes | 3,266 | ceterms:estimatedDuration |
| Where the program is offered | 3,266 | ceterms:availableAt |
| What kind of provider it is | 3,266 | ceterms:agentSectorType |
| What it costs a student funded under WIOA | 3,266 | ceterms:CostProfile with ceterms:directCostType |
| The state's ten-year outlook for the occupation | 3,250 | None used |
| Four of the nine reported outcome measures | 2,099 | qdata:Metric / qdata:Observation |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
- The CIP code for the field of study. CTDL has an instructional-program property that takes a CIP alignment, in the same shape this export already uses for the occupation's SOC code. The CIP code the source filed is not carried.
- Online, in person, or both. CTDL has a delivery-type property, but its value has to be a concept from a controlled vocabulary that credreg.net serves as a web page rather than as data. This export emits no concept it cannot check against machine-readable data, so the format is not carried.
- How long the program takes. CTDL has a property for the estimated duration of a learning opportunity. The source's length in weeks and hours is not carried, and neither is the competency-based flag, which means a program finishes when the student can do the work and so has no fixed length by design.
- Where the program is offered. CTDL has a property for where a learning opportunity is available. The program's location, and the region this project derives from it, are not carried. For a separate reason, no address is put on the organization either: the location on a record is the program's, not necessarily the provider's.
- What kind of provider it is. The source's provider category does not map onto CTDL's agent-sector vocabulary without judgment calls, and that vocabulary is served as a web page rather than as data. The organization carries the name the source filed and nothing else.
- What it costs a student funded under WIOA. That is a different cost to a different payer, and CTDL can carry it as a second cost profile distinguished by a concept from a vocabulary served as a web page rather than as data. Only the out-of-pocket total is carried.
- The state's ten-year outlook for the occupation. This export projects the federal training record. California's projections for the occupation each program feeds — median wage, expected openings, growth — are joined to the program everywhere else on this site and are not carried here. They describe an occupation rather than this program, and hanging them off the program would assert that the program leads to that wage, which the source does not say. The occupation code itself is carried, so the alignment is stated and the projection is not.
- Four of the nine reported outcome measures. The source reports nine WIOA performance measures and this export projects five. Total served, total exited, total completed and employment in the fourth quarter after exit are reported and are not carried. The statistics layer could express them in exactly the same shape as the five that are.
Separately, 1 program(s) report a cost total that a suppressed component makes a floor rather than a price. CTDL's price property has no way to say "at least", so no cost is published for those rather than publishing a floor as though it were the fee.
What a separate validator found
The export checks itself, but every one of those checks was written by the same hand as the export, against the same reading of the same schema — which is the reading a mistake would survive. So it is also run through ctdl-validate 0.1.0, a separate tool with its own copies of Credential Engine's schema and a citation for every rule it applies. It has the same author as this site, so it is a second reading of the schema and not an outside review. It makes no network request and submits nothing anywhere.
5,907 entities were checked.
| Severity | Findings |
|---|---|
| Error (blocking) | 0 |
| Warning | 5,907 |
| Information | 0 |
| Unverifiable | 0 |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
No errors, and one warning, on every entity in the graph. The warning is the tension this export already had on the record: the published grammar says an identifier is a random UUID, and an export that has to produce the same identifiers every time cannot use a random one. Nothing else came back — no property used on a class that does not declare it, no reference pointing at nothing, no relationship stated in one direction and contradicted in the other, no invented term.
| Finding | Severity | Times | Entities | State |
|---|---|---|---|---|
CTID_NOT_UUIDV4 | Warning | 5,907 | 5,907 | Accepted, with a reason |
Counted from the export of the 2026-08-07 dataset snapshot, at the moment that export ran.
An accepted finding is a decision on the record, not a filter: it stays counted here and in the machine-readable statement. Any finding whose code has not been reasoned about fails the export instead of appearing quietly among the others.
What the validator could and could not judge
A clean result is only as wide as the vocabulary the checker holds. This one drives its structural checks from the core schema documents it carries, so it was in a position to judge 4 of the 7 classes and 17 of the 24 properties this export emits.
The rest are the outcome-statistics layer, which publishes its own schema document that the validator does not carry, plus one currency property. A term a checker has never heard of is one it declines to judge, not one it approves.
- Classes not judged
qdata:DataSetProfile, qdata:Metric, qdata:Observation- Properties not judged
qdata:hasMetric, qdata:hasObservation, qdata:isObservationOf, qdata:median, qdata:metricType, qdata:relevantDataSetFor, schema:currency
Those terms were checked by the export itself against the statistics schema, fetched and recorded in this project's provenance. That is a weaker guarantee than an outside opinion, and it is named as one.
What each mapping rests on
Every class and property here was chosen against a published definition rather than from memory, and the export refuses at build time to emit a term the vocabulary does not define. These are the primary definitions, not summaries of them.
Credential Engine publishes these in English only, so they are marked as English on this page.
- ceterms:LearningProgram — "Set of learning opportunities that leads to an outcome, usually a credential like a degree or certificate". Every record here is a state-listed training program; ceterms:Course is for a single structured sequence, which the source does not distinguish, so it is never used
- ceterms:offeredBy — "Agent that offers the resource". Chosen over ceterms:ownedBy, whose definition is an enforceable claim or legal title: the training list asserts that a provider offers a program and says nothing about title
- ceterms:occupationType — a credential alignment whose usage note names SOC among the expected frameworks. The alignment carries the code the source filed, and a title only where this dataset matched that exact code
- ceterms:estimatedCost — a cost profile with a price and a currency. Emitted only where the source total is complete, because a total with a suppressed component is a floor and the price property cannot say "at least"
- qdata:relevantDataSet — "Data Set on which earnings or employment data is based", which names ceterms:LearningProgram in its own domain rather than relying on a subclass relation
- qdata:Metric and qdata:Observation — what is being measured, and the value observed for it. Counts carry a value, earnings carry a median with a currency, and rates carry a percentage, which is why the source's 0–1 fractions are multiplied by 100
- About the CTID — "Each CTID is made up of a standard UUID v4 prefixed with ce-". This export uses a v5 so that re-exporting the same data reproduces the same identifiers, which is the one thing a v4 cannot do, and publishes the warning that results
- Schema-Development issue #1080, filed from this project: ceterms:aggregateData did not list ceterms:LearningProgram. The maintainers' answer settled the design — the Registry no longer accepts aggregateData for publishing, and the supported pattern is the data-set profile this export now uses
Read the export, including the reason recorded beside every mapping →
Getting the export, and rebuilding it
The graph is about 17 MB of JSON-LD, which is too large to commit and too specific to one snapshot to serve as though it were current. It is built on demand and packaged with a checksum, and the two statements this page renders from are published beside it.
The statements this page is built from, as published:
- Coverage statement (what is carried, and what is not)
- Validation statement (what the validator found)
To rebuild the graph from source: clone the repository, run the pipeline to fetch the public federal and state data, then run the export and the validation. Both are single commands and both are deterministic — the same dataset always produces byte-identical output, so a rebuild can be compared against a published one directly.
Packaged exports are published as releases: https://github.com/ChelseaKR/afterward/releases
Corrections
If a mapping here is wrong, or a property is being used in a way the schema does not intend, that is worth an issue. This is a demonstration and the point of it is to be checkable; a correction from somebody who works on this vocabulary is the most useful thing this page could produce.