Experimental vendor proposal · opt-in publication · operator-managed client integration
Purpose
/llms-tldr.txt is a bounded, deterministic publication of literal stored
WordPress content. It exists for an operator who wants one inspectable site
briefing with an explicit approximate token budget.
This CYBERMAPS-specific format complements llms.txt with a compact extract of stored content. It uses deterministic selection rather than AI generation. Operators can supply its URL to a compatible retrieval client or use the output in their own workflow.
Design constraints
The generator must:
- use the same eligible-content inventory as other Cybermaps publications;
- execute no shortcode, dynamic block, model, embedding, or remote service;
- state its exact stable selection order;
- account for every emitted byte, including the header;
- include only complete entries;
- report selected and omitted counts truthfully; and
- stay within the configured approximate budget.
Input inventory
Cybermaps\Discovery\PublicationInventory supplies published, eligible content
using the unified publication-eligibility decision. The inventory respects
configured AI post types, exclusions, redirects, off-resource canonicals,
Genesis/Mai settings, and supported SEO-plugin signals.
The generator does not inspect drafts, private resources, password-protected resources, or attachments.
Selection order
The stable order is:
- configured pinned resource IDs, in the operator’s order;
- configured content-type order;
- stored modified time descending; and
- WordPress object ID ascending as the final tie-breaker.
For each resource the generator renders a complete candidate entry. If adding that entry and the updated coverage header would exceed the budget, that entry is omitted and later candidates are still considered.
There is no quality score, taxonomy cluster, entity extractor, semantic deduplication, information-gain calculation, or inferred priority.
Extraction
Cybermaps\Content\VisibleTextExtractor derives literal visible text from
stored post content. It strips markup and ignores non-rendered stored elements;
it does not execute shortcodes or server-rendered dynamic blocks.
Each briefing entry contains:
- title;
- public URL;
- WordPress content type;
- stored last-modified timestamp when valid; and
- a fixed-length literal extract.
The extract may be incomplete by design, but an entry is never cut midway by the budgeter.
Budget calculation
The estimator is intentionally simple and disclosed:
estimated_tokens = ceil(number_of_UTF-8_bytes / 4)
The configured range is 1,000–200,000 estimated tokens; the default is 80,000. This is not a tokenizer for any particular model. It is a repeatable planning estimate.
Every calculation uses the complete candidate output:
- format metadata;
- selection and extraction disclosures;
- configured license assertion;
- coverage counts; and
- selected entries.
Therefore token_estimate <= token_budget for the emitted publication under the
documented estimator.
Header
The publication states:
Profile: Cybermaps Budgeted Site Briefing
Format-Version: 0.2-draft
Status: Experimental vendor proposal
Known-Automatic-Consumers: none documented
Selection-Method: ...
Content-Extraction: ...
Token-Estimate-Method: ...
Token-Budget: ...
Coverage: selected N of E eligible; omitted O; partial 0
partial is always zero in 0.2 because the generator includes or omits whole
entries.
A license value is labeled License-Assertion because it is supplied by the
site operator. Cybermaps does not verify ownership, legal scope, or
enforceability.
Caching and delivery
The canonical dynamic response uses a one-hour WordPress transient. Content and
relevant settings changes invalidate the discovery cache. With full static
publication enabled, the same generated body can be materialized at
/llms-tldr.txt; localized variants can also be materialized when translation
output is configured.
Physical delivery bypasses PHP. It therefore cannot appear in PHP-observed crawler analytics and uses server-controlled response headers.
Limitations
- The byte-to-token ratio is approximate and model-independent.
- Stored-text extraction cannot represent runtime output from dynamic blocks or shortcodes.
- The fixed extract can omit important context.
- Pinned IDs express operator priority, not objective importance.
- Stable modified-date ordering favors recent edits, not quality.
- No semantic duplicate detection occurs.
- Client integration is configured by the operator.
- A static file remains stale until the next complete reconciliation.
Appropriate use
An operator can inspect or deliver the briefing as one bounded snapshot alongside
the canonical pages and /llms.txt. Any client should treat each URL as the
authoritative source for full context.
The draft is opt-in. Future format changes require a version increment, fixtures for exact budget behavior, and migration documentation so integrations can follow a defined upgrade path.