Skip to content

calternal_notes_core::nlp::runs

Composer highlight runs — UTF-16 spans for the mirror-overlay.

The composer highlights recognized tokens (dates, times, tags, …) as the user types. To GUARANTEE the highlight spans never disagree with the parsed values, the spans are derived from the SAME consumed token set the parser builds — NOT re-parsed from the text. We re-run the existing recognizer pipeline and attribute each consumed token to the recognizer/stage that claimed it.

KEEP IN LOCK-STEP (this is the whole contract)

Section titled “KEEP IN LOCK-STEP (this is the whole contract)”

This module RE-WALKS the production pipelines:

  • the ENTRY path mirrors crate::nlp::parse_internal (mod.rs) — the recognizer call ORDER and each consume-loop (including the from-prefix and standalone-duration double-claim guards) are replicated here so our consumed set is byte-for-byte identical.
  • the TASK path calls the SAME stage functions via crate::nlp::tasks::run_task_stages’s component recognizers, snapshotting which token indices each detect_* call newly claims. As of #27 T8 the time-of-day and landing-aware-tag stages call tasks::attach_time_of_day / tasks::consume_landing_tags DIRECTLY (not a re-derived copy) — those two stages can no longer drift; only the recurrence/priority/date-role stages (thin wrappers over classifier::detect_*) are still replicated call-by-call here.

If the recognizer order or any consume logic changes in parse_internal / tasks::run_task_stages, this module MUST be updated in step. The run_coverage_equals_pipeline_consumed_set test is the drift guard for the task path: it asserts the union of token indices our emitted runs cover equals the pipeline’s final consumed set — no more, no less.

  1. Tokenize (same tokenizer the pipelines use); resolve the reference date the same way (civil::parse_reference_date).
  2. Re-run the recognizers in order; label each newly-consumed token index with the claiming stage’s RunKind.
  3. Map each labelled token’s BYTE span to UTF-16 code-unit offsets via ONE left-to-right walk of the text (byte_to_utf16_map) — JS/Swift string indices are UTF-16, and a leading multi-byte char must not shift the spans.
  4. Merge adjacent same-kind phrase runs separated only by ASCII whitespace into one run (e.g. every 2 weeks → one Recurrence, Jan 15 → one Date). Tags deliberately do NOT merge: each tag is one editable chip.
  5. Sort by start.

Source: crates/calternal-notes-core/src/nlp/runs.rs

pub struct Run

A highlight run: a half-open [start, end) range in UTF-16 code units (the unit JS/Swift strings index by) and the RunKind it represents. camelCase keeps the field names idiomatic on the wire (cf. TaskRecurrence).

Fields

  • pub start: usize
  • pub end: usize
  • pub kind: RunKind

Implements: Clone, Copy, PartialEq, Eq, Debug

Source: crates/calternal-notes-core/src/nlp/runs.rs:65

pub enum RunKind

The kind of recognized token a highlight run covers. Serialized lowercase to match the JS wire form the WASM layer consumes (cf. TaskPriority).

Variants

  • Date
  • Time
  • Duration
  • Priority
  • Recurrence
  • Tag
  • Anchor

Implements: Clone, Copy, PartialEq, Eq, Debug

Source: crates/calternal-notes-core/src/nlp/runs.rs:51

pub fn parse_entry_runs(text: &str, now: &str) -> Vec<Run>

Derive highlight runs for the ENTRY (journal) composer pipeline.

Source: crates/calternal-notes-core/src/nlp/runs.rs:82

pub fn parse_task_runs(text: &str, now: &str) -> Vec<Run>

Derive highlight runs for the TASK composer pipeline.

Source: crates/calternal-notes-core/src/nlp/runs.rs:87