Data provenance guide
Trace tokenized asset data from source to model output
A model can organize descriptions and extract claims from tokenized asset records, but useful output begins with reliable identity and traceable evidence. Separate ledger observations, external documents, numerical features and model conclusions before joining them into a pipeline. This guide focuses on data engineering: which records to retain, how to describe their origin and how to evaluate derived labels. It also keeps blockchain tokens distinct from language model tokens, whose counts and identifiers belong to a different system.
Identify the subject and the observed state
Use the strongest identifier available for each source. A chain identifier and contract or mint address can distinguish ledger assets more reliably than a display symbol alone. A document needs its own source location, observation time and version or content hash. Preserve both when an external description refers to a ledger record.
Distinguish the time your service collected a document from the period its contents describe. Record the relevant block reference for a ledger observation. Keep missing metadata explicit instead of guessing a value that later appears authoritative. These decisions let a reviewer reconstruct which state the model actually saw. They also make updates easier to reconcile when a source changes, a document is corrected or two records turn out to describe different subjects.
Keep derivation records beside numerical features
W3C PROV provides a vocabulary for describing sources, processing activities and responsible actors. Apply that discipline to each stage that changes the meaning of data.
- Retain a source reference and content hash for captured evidence.
- Record parser and normalization versions for deterministic transformations.
- Describe each feature’s definition, units and missing-value policy.
- Preserve row identifiers and feature order when values become a tensor.
- Record the model, prompt and label taxonomy used for derived results.
A numerical array is convenient for computation, but it does not explain itself. Keep observed values distinguishable from model-generated labels. Changing a feature’s definition can change the prediction contract even when its element type and tensor shape remain identical.
Classify evidence within an explicit scope
Choose questions the supplied records can answer. A classifier can identify whether a passage describes a reporting process or categorize a document under a reviewed taxonomy. It should preserve the difference between a claim made by a source and a conclusion independently established by the application. Allow an unknown result when the evidence is insufficient.
Require extracted fields to reference supplied passages. Evaluate with missing, conflicting, duplicated and outdated records, and keep near-duplicates out of opposing evaluation partitions. When a source is corrected, mark earlier observations as superseded and recompute affected features. Return evidence references and result versions so callers can inspect changes. This workflow helps organize information; it does not turn a model label into a financial recommendation or a guarantee about an asset.
Use the official reference for the documented interface; the workflow recommendations above are Tensor API Lab guidance.