calternal_notes_core::nlp::runs
Composer highlight runs — UTF-16 spans for the mirror-overlay.
The composer highlights recognized tokens (dates, times, tags, …) as the user
types. To GUARANTEE the highlight spans never disagree with the parsed values,
the spans are derived from the SAME consumed token set the parser builds —
NOT re-parsed from the text. We re-run the existing recognizer pipeline and
attribute each consumed token to the recognizer/stage that claimed it.
KEEP IN LOCK-STEP (this is the whole contract)
Section titled “KEEP IN LOCK-STEP (this is the whole contract)”This module RE-WALKS the production pipelines:
- the ENTRY path mirrors
crate::nlp::parse_internal(mod.rs) — the recognizer call ORDER and each consume-loop (including thefrom-prefix and standalone-duration double-claim guards) are replicated here so ourconsumedset is byte-for-byte identical. - the TASK path calls the SAME stage functions via
crate::nlp::tasks::run_task_stages’s component recognizers, snapshotting which token indices eachdetect_*call newly claims. As of #27 T8 the time-of-day and landing-aware-tag stages calltasks::attach_time_of_day/tasks::consume_landing_tagsDIRECTLY (not a re-derived copy) — those two stages can no longer drift; only the recurrence/priority/date-role stages (thin wrappers overclassifier::detect_*) are still replicated call-by-call here.
If the recognizer order or any consume logic changes in parse_internal /
tasks::run_task_stages, this module MUST be updated in step. The
run_coverage_equals_pipeline_consumed_set test is the drift guard for the
task path: it asserts the union of token indices our emitted runs cover equals
the pipeline’s final consumed set — no more, no less.
How a span is built
Section titled “How a span is built”- Tokenize (same tokenizer the pipelines use); resolve the reference date the
same way (
civil::parse_reference_date). - Re-run the recognizers in order; label each newly-consumed token index with
the claiming stage’s
RunKind. - Map each labelled token’s BYTE span to UTF-16 code-unit offsets via ONE
left-to-right walk of the text (
byte_to_utf16_map) — JS/Swift string indices are UTF-16, and a leading multi-byte char must not shift the spans. - Merge adjacent same-kind phrase runs separated only by ASCII whitespace
into one run (e.g.
every 2 weeks→ oneRecurrence,Jan 15→ oneDate). Tags deliberately do NOT merge: each tag is one editable chip. - Sort by
start.
Source: crates/calternal-notes-core/src/nlp/runs.rs
Structs
Section titled “Structs”pub struct RunA highlight run: a half-open [start, end) range in UTF-16 code units
(the unit JS/Swift strings index by) and the RunKind it represents.
camelCase keeps the field names idiomatic on the wire (cf. TaskRecurrence).
Fields
pub start: usizepub end: usizepub kind: RunKind
Implements: Clone, Copy, PartialEq, Eq, Debug
Source: crates/calternal-notes-core/src/nlp/runs.rs:65
RunKind
Section titled “RunKind”pub enum RunKindThe kind of recognized token a highlight run covers. Serialized lowercase to
match the JS wire form the WASM layer consumes (cf. TaskPriority).
Variants
DateTimeDurationPriorityRecurrenceTagAnchor
Implements: Clone, Copy, PartialEq, Eq, Debug
Source: crates/calternal-notes-core/src/nlp/runs.rs:51
Functions
Section titled “Functions”parse_entry_runs
Section titled “parse_entry_runs”pub fn parse_entry_runs(text: &str, now: &str) -> Vec<Run>Derive highlight runs for the ENTRY (journal) composer pipeline.
Source: crates/calternal-notes-core/src/nlp/runs.rs:82
parse_task_runs
Section titled “parse_task_runs”pub fn parse_task_runs(text: &str, now: &str) -> Vec<Run>Derive highlight runs for the TASK composer pipeline.