Skip to content

calternal-search

Rebuildable Tantivy search index for calternal

Rebuildable search index for files, Notes, and plain text.

The Index stores only derived data. The Data directory remains the source of truth, and every scan reads through calternal-fs directory handles.

Source: crates/calternal-search/src/lib.rs

pub struct ContextCitation

A source that an Ask answer can cite and the client can open by deep link.

Fields

  • pub marker: String: Short marker required in the streamed answer, such as [C1].
  • pub plugin_id: String: Plugin that owns the search result.
  • pub item_id: String: Stable result identifier supplied by that Plugin’s search provider.
  • pub kind: Option<String>: Search result kind, when the provider knows it.
  • pub title: String: Main result label.
  • pub href: String: Client route that opens the result.

Implements: Clone, Debug, Deserialize, Eq, PartialEq, Serialize

Source: crates/calternal-search/src/retrieval.rs:31

pub struct Indexer

No doc comment.

Implements: Clone

pub async fn start(
root: Root,
user_data_directory: impl AsRef<Path>,
pool: Pool<Sqlite>,
) -> Result<Self, IndexerError>

Open the Index and start watching user_data_directory/users recursively. The first reconciliation runs in the background after the watcher starts.

pub async fn start_with_pdf_extractor(
root: Root,
user_data_directory: impl AsRef<Path>,
pool: Pool<Sqlite>,
pdf_extractor: Arc<dyn PdfTextExtractor>,
) -> Result<Self, IndexerError>

Start watching Homes and extract PDF text through the selected worker.

pub async fn start_deferred(
root: Root,
user_data_directory: impl AsRef<Path>,
pool: Pool<Sqlite>,
) -> Result<Self, IndexerError>

Start watching Homes and defer the first scan until the server opens its listener. The persisted Index stays available to queries meanwhile.

Indexer::start_deferred_with_pdf_extractor

Section titled “Indexer::start_deferred_with_pdf_extractor”
pub async fn start_deferred_with_pdf_extractor(
root: Root,
user_data_directory: impl AsRef<Path>,
pool: Pool<Sqlite>,
pdf_extractor: Arc<dyn PdfTextExtractor>,
) -> Result<Self, IndexerError>

Deferred-start variant that uses the selected PDF text worker.

Indexer::start_deferred_with_pools_and_pdf_extractor

Section titled “Indexer::start_deferred_with_pools_and_pdf_extractor”
pub async fn start_deferred_with_pools_and_pdf_extractor(
root: Root,
user_data_directory: impl AsRef<Path>,
writer_pool: Pool<Sqlite>,
reader_pool: Pool<Sqlite>,
pdf_extractor: Arc<dyn PdfTextExtractor>,
) -> Result<Self, IndexerError>

Deferred-start variant with separate database read and write pools.

Search frecency reads use the read-only pool so indexing, imports and other writes on the single writer pool do not delay keyword results.

pub async fn open(root: Root, pool: Pool<Sqlite>) -> Result<Self, IndexerError>

Open an Index without a watcher or initial scan.

The server uses this for the one-shot reindex command. Call Indexer::rebuild before the process exits.

pub async fn open_with_pdf_extractor(
root: Root,
pool: Pool<Sqlite>,
pdf_extractor: Arc<dyn PdfTextExtractor>,
) -> Result<Self, IndexerError>

Open an Index without a watcher or initial scan and select PDF parsing.

pub fn notify_changed(&self, path: RelPath) -> Result<(), IndexerError>

Add or refresh a file or folder after a server-side write.

The path must be below a Home and outside Trash and .calternal data. Filesystem access later happens through calternal-fs.

pub fn notify_removed(&self, path: RelPath) -> Result<(), IndexerError>

Remove a file or folder after a server-side delete or move.

pub async fn reconcile(&self) -> Result<(), IndexerError>

Reconcile all Homes against the filesystem and SQLite manifest.

pub async fn reconcile_at_start(&self) -> Result<(), IndexerError>

Start the first full scan after the server has opened its listener. The call completes when the background actor completes that scan.

pub async fn rebuild(&self) -> Result<(), IndexerError>

Clear the derived index and rebuild every document from the Data directory.

pub async fn rebuild_user(&self, user_id: &str) -> Result<(), IndexerError>

Build one complete private generation from the Home and incoming Shares while the migration generation continues serving this User.

pub async fn check_and_repair(&self) -> Result<IntegrityReport, IndexerError>

Compare the live Tantivy documents and SQLite manifest with the filesystem, then repair missing, stale, or duplicate documents.

pub fn integrity_status(&self) -> IntegrityReport

Return the latest completed integrity check and whether one is running.

pub fn indexing_status(&self) -> SearchIndexStatus

Return progress for the current full scan or repair pass.

pub async fn shutdown(&self) -> Result<(), IndexerError>

Stop the Search actor at a safe batch boundary and wait no longer than the server stop budget. An interrupted repair is rebuilt on next start.

pub async fn search_grouped(
&self,
context: &calternal_plugin::SearchContext,
query: &str,
limit: usize,
) -> Result<Vec<SearchGroup>, IndexerError>

Search the index and return results grouped by Notes, files, and folders.

pub async fn annotate_broken_link_hits(
&self,
user_id: &str,
hits: &mut [SearchHit],
) -> Result<(), IndexerError>

Attach the first unresolved link to each matching Note result. Search flags select the Notes in Tantivy; this small batched lookup supplies the source text and line needed by the preview and exact jump.

pub async fn retain_live(&self, hits: Vec<SearchHit>) -> (Vec<SearchHit>, Vec<RelPath>)

Drop hits whose file no longer exists, and queue their removal from this Index. Returns the kept hits (in order) and the vanished Index paths, so a caller with a second index (semantic) can queue their removal there too.

A hit’s Index path is its id; a log entry’s id adds a #^block or #log-<n> fragment to its Daily note path. File names may contain #, so only log hits are split.

pub async fn record_opened(
&self,
user_id: &str,
plugin_id: &str,
result_id: &str,
) -> Result<(), IndexerError>

Record one authenticated result open for per-User ranking.

Source: crates/calternal-search/src/indexer.rs:228

pub struct InProcessPdfTextExtractor;

Extract text in a resource-limited caller process, or in small fixture tests.

Implements: Clone, Copy, Debug, Default, PdfTextExtractor

Source: crates/calternal-search/src/pdf.rs:32

pub struct IntegrityReport

Latest full comparison between the filesystem and the derived search Index.

Fields

  • pub checked_items: u64: Number of searchable filesystem paths checked, including folders.
  • pub missing_items: u64: Documents missing from the Index before repair.
  • pub stale_items: u64: Documents absent from the filesystem or duplicated in the Index.
  • pub manifest_mismatches: u64: Filesystem paths whose SQLite manifest row needed an update.
  • pub repaired_items: u64: Source paths changed by the repair.
  • pub checked_at_unix_ms: Option<i64>: Unix time in milliseconds for the last completed check.
  • pub running: bool: True while the actor is scanning and repairing the Index.
  • pub healthy: bool: True when the post-repair Index matches the filesystem and manifest.
  • pub last_error: Option<String>: The latest scan error, if a check did not complete.

Implements: Clone, Debug, Default, PartialEq, Eq, Serialize, ToSchema

Source: crates/calternal-search/src/indexer.rs:206

pub struct RetrievalContext

Search results formatted for an Agent prompt and a citations SSE event.

Fields

  • pub prompt_context: String: Bounded JSON Lines. Retrieved values remain data and cannot add prompt sections or citation markers by injecting quotes or newlines.
  • pub citations: Vec<ContextCitation>: Stable result identities paired with client deep links.

Implements: Clone, Debug, Default, Deserialize, Eq, PartialEq, Serialize

Source: crates/calternal-search/src/retrieval.rs:48

pub struct SearchGroup

No doc comment.

Fields

  • pub kind: SearchGroupKind
  • pub hits: Vec<calternal_api::SearchHit>

Implements: Clone, Debug, PartialEq

Source: crates/calternal-search/src/query.rs:42

pub struct SearchPlugin

No doc comment.

Implements: Plugin

pub fn new(indexer: Indexer) -> Self

Attach the existing Indexer to the Search Plugin (#911; DESIGN §8). The cloneable handle shares the same indexing service; this does not start a second writer or change the Index generation.

pub fn with_semantic(indexer: Indexer, semantic: SemanticIndexer) -> Self

Construct the live hybrid provider. The semantic index remains optional at the data level because model download and embedding work are bounded background tasks.

pub fn metadata_only() -> Self

Construct the plugin manifest without a live Indexer.

The server uses this while generating API metadata. Live request registries must use SearchPlugin::new or SearchPlugin::with_semantic.

Source: crates/calternal-search/src/plugin.rs:13

pub struct SubprocessPdfTextExtractor

Run the parser in a child process and kill it when the deadline expires.

Implements: Clone, Debug, PdfTextExtractor

SubprocessPdfTextExtractor::current_executable

Section titled “SubprocessPdfTextExtractor::current_executable”
pub fn current_executable() -> std::io::Result<Self>

Use the running server executable for the hidden PDF extraction command.

pub fn new(executable: PathBuf, timeout: Duration) -> Self

Construct an extractor for a server executable or a test helper.

Source: crates/calternal-search/src/pdf.rs:42

pub enum IndexerError

No doc comment.

Variants

  • Filesystem(#[from] FsError)
  • Database(#[from] sqlx::Error)
  • Tantivy(#[from] tantivy::TantivyError)
  • Query(#[from] tantivy::query::QueryParserError)
  • SearchQuery(#[from] crate::operators::SearchQueryError)
  • Io(#[from] std::io::Error)
  • IndexDirectory(String)
  • IndexMetadata(String)
  • Watcher(String)
  • InvalidRoots
  • WorkerStopped
  • WriterLockTimeout(u64)
  • ShutdownRequested
  • ShutdownTimeout
  • QueueFull
  • InvalidOpenRecord
  • Background(String)
  • Thread(std::io::Error)
  • BlockingTask(tokio::task::JoinError)

Implements: Debug, Error

Source: crates/calternal-search/src/indexer.rs:163

pub enum SearchGroupKind

No doc comment.

Variants

  • Notes
  • Files
  • Folders

Implements: Clone, Copy, Debug, Eq, Ord, PartialEq, PartialOrd

Source: crates/calternal-search/src/query.rs:35

pub enum SearchQueryError

No doc comment.

Variants

  • TooLong
  • TermTooLong
  • ControlCharacter
  • UnfinishedQuote
  • TooManyParts
  • MissingValue
  • InvalidValue

Implements: Clone, Copy, Debug, PartialEq, Eq, Error

Source: crates/calternal-search/src/operators.rs:20

pub trait PdfTextExtractor: Send + Sync + 'static

Extract searchable text from one PDF file.

fn extract(&self, bytes: &[u8]) -> Option<String>;

Return extracted text, or None when the PDF is invalid or over a cap.

Source: crates/calternal-search/src/pdf.rs:25

pub fn apply_pdf_extraction_limits() -> std::io::Result<()>

Apply the same hard boundary to direct hidden-command calls (#781). Call this before input allocation or Tokio startup. Existing stricter limits stay in force. Decoded stream buffers share this aggregate cap.

Source: crates/calternal-search/src/pdf.rs:127

pub fn extract_pdf_text(bytes: &[u8]) -> Option<String>

Extract text from at most 128 pages and 256 KiB of output.

The caller controls the hard deadline. This function also checks elapsed time between pages to avoid doing more work after a slow page.

Source: crates/calternal-search/src/pdf.rs:182

pub fn is_searchable_path(path: &RelPath) -> bool

Return whether a path belongs in the search Index.

Every caller must use this policy before queuing watcher or filesystem changes. Hidden and calternal-owned paths stay out of keyword and semantic indexes; the shared path policy keeps this entry point aligned with the other activity surfaces.

Source: crates/calternal-search/src/indexer.rs:4043

pub fn migration_set() -> MigrationSet

Return the migration that stores the rebuildable file manifest in the Index.

Source: crates/calternal-search/src/lib.rs:38

pub fn parse_search_query(input: &str) -> Result<ParsedSearchQuery, SearchQueryError>

Parse text and search operators into the stable API query shape.

Quoted phrases remain one term. Operators may appear in any order and may repeat. Unknown name:value tokens stay plain text so new operators do not make older clients lose searchable words.

Source: crates/calternal-search/src/operators.rs:42

pub fn retrieval_query(question: &str) -> String

Remove common question wording so keyword retrieval can match the subject.

The semantic provider still receives these subject terms. Keeping question words out avoids requiring a note to contain words such as “when” or “spent” before its useful names and places can match.

Source: crates/calternal-search/src/retrieval.rs:199

pub fn retrieve_for_context(
results: &[SearchHit],
max_sources: usize,
max_context_bytes: usize,
) -> RetrievalContext

Convert visible search results into a bounded Agent context.

Search already applies the authenticated User’s allowed roots. This API keeps that result set intact, rejects non-local links, removes duplicates, and applies an independent byte limit before the context reaches an Agent.

Source: crates/calternal-search/src/retrieval.rs:246

pub async fn stabilize_for_context(
db: &Db,
user_id: &str,
results: &[SearchHit],
) -> Vec<SearchHit>

Replace path-based search result ids with current stable object identities.

The search response has already passed through each provider’s visibility checks. This step only resolves those returned paths through the Files and Notes indexes; it does not read the Data directory or broaden access.

Source: crates/calternal-search/src/retrieval.rs:61

pub const MAX_PDF_BYTES: usize

No doc comment.

Source: crates/calternal-search/src/pdf.rs:18

pub const MAX_PDF_TEXT_BYTES: usize

No doc comment.

Source: crates/calternal-search/src/pdf.rs:19

pub const MAX_SAVED_SEARCH_NAME_CHARS: usize

Longest display name, in characters.

Source: crates/calternal-search/src/saved.rs:45

pub const MAX_SAVED_SEARCH_QUERY_CHARS: usize

Longest query, in characters (the search grammar’s own limit).

Source: crates/calternal-search/src/saved.rs:47

pub const MAX_SAVED_SEARCHES: usize

Most saved searches one user may keep.

Source: crates/calternal-search/src/saved.rs:49

pub const MAX_TEXT_BYTES: u64

No doc comment.

Source: crates/calternal-search/src/indexer.rs:52

pub const PDF_EXTRACT_COMMAND: &str

No doc comment.

Source: crates/calternal-search/src/pdf.rs:22

pub const SAVED_SEARCH_SCHEMA: &str

Schema tag written into every file (DESIGN §4: calternal.<kind>/<n>).

Source: crates/calternal-search/src/saved.rs:43