calternal-search
Rebuildable Tantivy search index for calternal
Rebuildable search index for files, Notes, and plain text.
The Index stores only derived data. The Data directory remains the source
of truth, and every scan reads through calternal-fs directory handles.
Source: crates/calternal-search/src/lib.rs
Re-exports
Section titled “Re-exports”Structs
Section titled “Structs”ContextCitation
Section titled “ContextCitation”pub struct ContextCitationA source that an Ask answer can cite and the client can open by deep link.
Fields
pub marker: String: Short marker required in the streamed answer, such as[C1].pub plugin_id: String: Plugin that owns the search result.pub item_id: String: Stable result identifier supplied by that Plugin’s search provider.pub kind: Option<String>: Search result kind, when the provider knows it.pub title: String: Main result label.pub href: String: Client route that opens the result.
Implements: Clone, Debug, Deserialize, Eq, PartialEq, Serialize
Source: crates/calternal-search/src/retrieval.rs:31
Indexer
Section titled “Indexer”pub struct IndexerNo doc comment.
Implements: Clone
Indexer::start
Section titled “Indexer::start”pub async fn start( root: Root, user_data_directory: impl AsRef<Path>, pool: Pool<Sqlite>, ) -> Result<Self, IndexerError>Open the Index and start watching user_data_directory/users recursively.
The first reconciliation runs in the background after the watcher starts.
Indexer::start_with_pdf_extractor
Section titled “Indexer::start_with_pdf_extractor”pub async fn start_with_pdf_extractor( root: Root, user_data_directory: impl AsRef<Path>, pool: Pool<Sqlite>, pdf_extractor: Arc<dyn PdfTextExtractor>, ) -> Result<Self, IndexerError>Start watching Homes and extract PDF text through the selected worker.
Indexer::start_deferred
Section titled “Indexer::start_deferred”pub async fn start_deferred( root: Root, user_data_directory: impl AsRef<Path>, pool: Pool<Sqlite>, ) -> Result<Self, IndexerError>Start watching Homes and defer the first scan until the server opens its listener. The persisted Index stays available to queries meanwhile.
Indexer::start_deferred_with_pdf_extractor
Section titled “Indexer::start_deferred_with_pdf_extractor”pub async fn start_deferred_with_pdf_extractor( root: Root, user_data_directory: impl AsRef<Path>, pool: Pool<Sqlite>, pdf_extractor: Arc<dyn PdfTextExtractor>, ) -> Result<Self, IndexerError>Deferred-start variant that uses the selected PDF text worker.
Indexer::start_deferred_with_pools_and_pdf_extractor
Section titled “Indexer::start_deferred_with_pools_and_pdf_extractor”pub async fn start_deferred_with_pools_and_pdf_extractor( root: Root, user_data_directory: impl AsRef<Path>, writer_pool: Pool<Sqlite>, reader_pool: Pool<Sqlite>, pdf_extractor: Arc<dyn PdfTextExtractor>, ) -> Result<Self, IndexerError>Deferred-start variant with separate database read and write pools.
Search frecency reads use the read-only pool so indexing, imports and other writes on the single writer pool do not delay keyword results.
Indexer::open
Section titled “Indexer::open”pub async fn open(root: Root, pool: Pool<Sqlite>) -> Result<Self, IndexerError>Open an Index without a watcher or initial scan.
The server uses this for the one-shot reindex command. Call
Indexer::rebuild before the process exits.
Indexer::open_with_pdf_extractor
Section titled “Indexer::open_with_pdf_extractor”pub async fn open_with_pdf_extractor( root: Root, pool: Pool<Sqlite>, pdf_extractor: Arc<dyn PdfTextExtractor>, ) -> Result<Self, IndexerError>Open an Index without a watcher or initial scan and select PDF parsing.
Indexer::notify_changed
Section titled “Indexer::notify_changed”pub fn notify_changed(&self, path: RelPath) -> Result<(), IndexerError>Add or refresh a file or folder after a server-side write.
The path must be below a Home and outside Trash and .calternal data.
Filesystem access later happens through calternal-fs.
Indexer::notify_removed
Section titled “Indexer::notify_removed”pub fn notify_removed(&self, path: RelPath) -> Result<(), IndexerError>Remove a file or folder after a server-side delete or move.
Indexer::reconcile
Section titled “Indexer::reconcile”pub async fn reconcile(&self) -> Result<(), IndexerError>Reconcile all Homes against the filesystem and SQLite manifest.
Indexer::reconcile_at_start
Section titled “Indexer::reconcile_at_start”pub async fn reconcile_at_start(&self) -> Result<(), IndexerError>Start the first full scan after the server has opened its listener. The call completes when the background actor completes that scan.
Indexer::rebuild
Section titled “Indexer::rebuild”pub async fn rebuild(&self) -> Result<(), IndexerError>Clear the derived index and rebuild every document from the Data directory.
Indexer::rebuild_user
Section titled “Indexer::rebuild_user”pub async fn rebuild_user(&self, user_id: &str) -> Result<(), IndexerError>Build one complete private generation from the Home and incoming Shares while the migration generation continues serving this User.
Indexer::check_and_repair
Section titled “Indexer::check_and_repair”pub async fn check_and_repair(&self) -> Result<IntegrityReport, IndexerError>Compare the live Tantivy documents and SQLite manifest with the filesystem, then repair missing, stale, or duplicate documents.
Indexer::integrity_status
Section titled “Indexer::integrity_status”pub fn integrity_status(&self) -> IntegrityReportReturn the latest completed integrity check and whether one is running.
Indexer::indexing_status
Section titled “Indexer::indexing_status”pub fn indexing_status(&self) -> SearchIndexStatusReturn progress for the current full scan or repair pass.
Indexer::shutdown
Section titled “Indexer::shutdown”pub async fn shutdown(&self) -> Result<(), IndexerError>Stop the Search actor at a safe batch boundary and wait no longer than the server stop budget. An interrupted repair is rebuilt on next start.
Indexer::search_grouped
Section titled “Indexer::search_grouped”pub async fn search_grouped( &self, context: &calternal_plugin::SearchContext, query: &str, limit: usize, ) -> Result<Vec<SearchGroup>, IndexerError>Search the index and return results grouped by Notes, files, and folders.
Indexer::annotate_broken_link_hits
Section titled “Indexer::annotate_broken_link_hits”pub async fn annotate_broken_link_hits( &self, user_id: &str, hits: &mut [SearchHit], ) -> Result<(), IndexerError>Attach the first unresolved link to each matching Note result. Search flags select the Notes in Tantivy; this small batched lookup supplies the source text and line needed by the preview and exact jump.
Indexer::retain_live
Section titled “Indexer::retain_live”pub async fn retain_live(&self, hits: Vec<SearchHit>) -> (Vec<SearchHit>, Vec<RelPath>)Drop hits whose file no longer exists, and queue their removal from this Index. Returns the kept hits (in order) and the vanished Index paths, so a caller with a second index (semantic) can queue their removal there too.
A hit’s Index path is its id; a log entry’s id adds a #^block or
#log-<n> fragment to its Daily note path. File names may contain
#, so only log hits are split.
Indexer::record_opened
Section titled “Indexer::record_opened”pub async fn record_opened( &self, user_id: &str, plugin_id: &str, result_id: &str, ) -> Result<(), IndexerError>Record one authenticated result open for per-User ranking.
Source: crates/calternal-search/src/indexer.rs:228
InProcessPdfTextExtractor
Section titled “InProcessPdfTextExtractor”pub struct InProcessPdfTextExtractor;Extract text in a resource-limited caller process, or in small fixture tests.
Implements: Clone, Copy, Debug, Default, PdfTextExtractor
Source: crates/calternal-search/src/pdf.rs:32
IntegrityReport
Section titled “IntegrityReport”pub struct IntegrityReportLatest full comparison between the filesystem and the derived search Index.
Fields
pub checked_items: u64: Number of searchable filesystem paths checked, including folders.pub missing_items: u64: Documents missing from the Index before repair.pub stale_items: u64: Documents absent from the filesystem or duplicated in the Index.pub manifest_mismatches: u64: Filesystem paths whose SQLite manifest row needed an update.pub repaired_items: u64: Source paths changed by the repair.pub checked_at_unix_ms: Option<i64>: Unix time in milliseconds for the last completed check.pub running: bool: True while the actor is scanning and repairing the Index.pub healthy: bool: True when the post-repair Index matches the filesystem and manifest.pub last_error: Option<String>: The latest scan error, if a check did not complete.
Implements: Clone, Debug, Default, PartialEq, Eq, Serialize, ToSchema
Source: crates/calternal-search/src/indexer.rs:206
RetrievalContext
Section titled “RetrievalContext”pub struct RetrievalContextSearch results formatted for an Agent prompt and a citations SSE event.
Fields
pub prompt_context: String: Bounded JSON Lines. Retrieved values remain data and cannot add prompt sections or citation markers by injecting quotes or newlines.pub citations: Vec<ContextCitation>: Stable result identities paired with client deep links.
Implements: Clone, Debug, Default, Deserialize, Eq, PartialEq, Serialize
Source: crates/calternal-search/src/retrieval.rs:48
SearchGroup
Section titled “SearchGroup”pub struct SearchGroupNo doc comment.
Fields
pub kind: SearchGroupKindpub hits: Vec<calternal_api::SearchHit>
Implements: Clone, Debug, PartialEq
Source: crates/calternal-search/src/query.rs:42
SearchPlugin
Section titled “SearchPlugin”pub struct SearchPluginNo doc comment.
Implements: Plugin
SearchPlugin::new
Section titled “SearchPlugin::new”pub fn new(indexer: Indexer) -> SelfAttach the existing Indexer to the Search Plugin (#911; DESIGN §8). The cloneable handle shares the same indexing service; this does not start a second writer or change the Index generation.
SearchPlugin::with_semantic
Section titled “SearchPlugin::with_semantic”pub fn with_semantic(indexer: Indexer, semantic: SemanticIndexer) -> SelfConstruct the live hybrid provider. The semantic index remains optional at the data level because model download and embedding work are bounded background tasks.
SearchPlugin::metadata_only
Section titled “SearchPlugin::metadata_only”pub fn metadata_only() -> SelfConstruct the plugin manifest without a live Indexer.
The server uses this while generating API metadata. Live request
registries must use SearchPlugin::new or SearchPlugin::with_semantic.
Source: crates/calternal-search/src/plugin.rs:13
SubprocessPdfTextExtractor
Section titled “SubprocessPdfTextExtractor”pub struct SubprocessPdfTextExtractorRun the parser in a child process and kill it when the deadline expires.
Implements: Clone, Debug, PdfTextExtractor
SubprocessPdfTextExtractor::current_executable
Section titled “SubprocessPdfTextExtractor::current_executable”pub fn current_executable() -> std::io::Result<Self>Use the running server executable for the hidden PDF extraction command.
SubprocessPdfTextExtractor::new
Section titled “SubprocessPdfTextExtractor::new”pub fn new(executable: PathBuf, timeout: Duration) -> SelfConstruct an extractor for a server executable or a test helper.
Source: crates/calternal-search/src/pdf.rs:42
IndexerError
Section titled “IndexerError”pub enum IndexerErrorNo doc comment.
Variants
Filesystem(#[from] FsError)Database(#[from] sqlx::Error)Tantivy(#[from] tantivy::TantivyError)Query(#[from] tantivy::query::QueryParserError)SearchQuery(#[from] crate::operators::SearchQueryError)Io(#[from] std::io::Error)IndexDirectory(String)IndexMetadata(String)Watcher(String)InvalidRootsWorkerStoppedWriterLockTimeout(u64)ShutdownRequestedShutdownTimeoutQueueFullInvalidOpenRecordBackground(String)Thread(std::io::Error)BlockingTask(tokio::task::JoinError)
Implements: Debug, Error
Source: crates/calternal-search/src/indexer.rs:163
SearchGroupKind
Section titled “SearchGroupKind”pub enum SearchGroupKindNo doc comment.
Variants
NotesFilesFolders
Implements: Clone, Copy, Debug, Eq, Ord, PartialEq, PartialOrd
Source: crates/calternal-search/src/query.rs:35
SearchQueryError
Section titled “SearchQueryError”pub enum SearchQueryErrorNo doc comment.
Variants
TooLongTermTooLongControlCharacterUnfinishedQuoteTooManyPartsMissingValueInvalidValue
Implements: Clone, Copy, Debug, PartialEq, Eq, Error
Source: crates/calternal-search/src/operators.rs:20
Traits
Section titled “Traits”PdfTextExtractor
Section titled “PdfTextExtractor”pub trait PdfTextExtractor: Send + Sync + 'staticExtract searchable text from one PDF file.
PdfTextExtractor::extract
Section titled “PdfTextExtractor::extract”fn extract(&self, bytes: &[u8]) -> Option<String>;Return extracted text, or None when the PDF is invalid or over a cap.
Source: crates/calternal-search/src/pdf.rs:25
Functions
Section titled “Functions”apply_pdf_extraction_limits
Section titled “apply_pdf_extraction_limits”pub fn apply_pdf_extraction_limits() -> std::io::Result<()>Apply the same hard boundary to direct hidden-command calls (#781). Call this before input allocation or Tokio startup. Existing stricter limits stay in force. Decoded stream buffers share this aggregate cap.
Source: crates/calternal-search/src/pdf.rs:127
extract_pdf_text
Section titled “extract_pdf_text”pub fn extract_pdf_text(bytes: &[u8]) -> Option<String>Extract text from at most 128 pages and 256 KiB of output.
The caller controls the hard deadline. This function also checks elapsed time between pages to avoid doing more work after a slow page.
Source: crates/calternal-search/src/pdf.rs:182
is_searchable_path
Section titled “is_searchable_path”pub fn is_searchable_path(path: &RelPath) -> boolReturn whether a path belongs in the search Index.
Every caller must use this policy before queuing watcher or filesystem changes. Hidden and calternal-owned paths stay out of keyword and semantic indexes; the shared path policy keeps this entry point aligned with the other activity surfaces.
Source: crates/calternal-search/src/indexer.rs:4043
migration_set
Section titled “migration_set”pub fn migration_set() -> MigrationSetReturn the migration that stores the rebuildable file manifest in the Index.
Source: crates/calternal-search/src/lib.rs:38
parse_search_query
Section titled “parse_search_query”pub fn parse_search_query(input: &str) -> Result<ParsedSearchQuery, SearchQueryError>Parse text and search operators into the stable API query shape.
Quoted phrases remain one term. Operators may appear in any order and may
repeat. Unknown name:value tokens stay plain text so new operators do not
make older clients lose searchable words.
Source: crates/calternal-search/src/operators.rs:42
retrieval_query
Section titled “retrieval_query”pub fn retrieval_query(question: &str) -> StringRemove common question wording so keyword retrieval can match the subject.
The semantic provider still receives these subject terms. Keeping question words out avoids requiring a note to contain words such as “when” or “spent” before its useful names and places can match.
Source: crates/calternal-search/src/retrieval.rs:199
retrieve_for_context
Section titled “retrieve_for_context”pub fn retrieve_for_context( results: &[SearchHit], max_sources: usize, max_context_bytes: usize,) -> RetrievalContextConvert visible search results into a bounded Agent context.
Search already applies the authenticated User’s allowed roots. This API keeps that result set intact, rejects non-local links, removes duplicates, and applies an independent byte limit before the context reaches an Agent.
Source: crates/calternal-search/src/retrieval.rs:246
stabilize_for_context
Section titled “stabilize_for_context”pub async fn stabilize_for_context( db: &Db, user_id: &str, results: &[SearchHit],) -> Vec<SearchHit>Replace path-based search result ids with current stable object identities.
The search response has already passed through each provider’s visibility checks. This step only resolves those returned paths through the Files and Notes indexes; it does not read the Data directory or broaden access.
Source: crates/calternal-search/src/retrieval.rs:61
Constants
Section titled “Constants”MAX_PDF_BYTES
Section titled “MAX_PDF_BYTES”pub const MAX_PDF_BYTES: usizeNo doc comment.
Source: crates/calternal-search/src/pdf.rs:18
MAX_PDF_TEXT_BYTES
Section titled “MAX_PDF_TEXT_BYTES”pub const MAX_PDF_TEXT_BYTES: usizeNo doc comment.
Source: crates/calternal-search/src/pdf.rs:19
MAX_SAVED_SEARCH_NAME_CHARS
Section titled “MAX_SAVED_SEARCH_NAME_CHARS”pub const MAX_SAVED_SEARCH_NAME_CHARS: usizeLongest display name, in characters.
Source: crates/calternal-search/src/saved.rs:45
MAX_SAVED_SEARCH_QUERY_CHARS
Section titled “MAX_SAVED_SEARCH_QUERY_CHARS”pub const MAX_SAVED_SEARCH_QUERY_CHARS: usizeLongest query, in characters (the search grammar’s own limit).
Source: crates/calternal-search/src/saved.rs:47
MAX_SAVED_SEARCHES
Section titled “MAX_SAVED_SEARCHES”pub const MAX_SAVED_SEARCHES: usizeMost saved searches one user may keep.
Source: crates/calternal-search/src/saved.rs:49
MAX_TEXT_BYTES
Section titled “MAX_TEXT_BYTES”pub const MAX_TEXT_BYTES: u64No doc comment.
Source: crates/calternal-search/src/indexer.rs:52
PDF_EXTRACT_COMMAND
Section titled “PDF_EXTRACT_COMMAND”pub const PDF_EXTRACT_COMMAND: &strNo doc comment.
Source: crates/calternal-search/src/pdf.rs:22
SAVED_SEARCH_SCHEMA
Section titled “SAVED_SEARCH_SCHEMA”pub const SAVED_SEARCH_SCHEMA: &strSchema tag written into every file (DESIGN §4: calternal.<kind>/<n>).