# GEO Module — `@sonordev/site-kit/llms`

Generative Engine Optimization (GEO) and Answer Engine Optimization (AEO) for Sonor-powered Next.js sites. Makes businesses visible to ChatGPT, Claude, Perplexity, Google AI Overviews, and voice assistants.

***

## Architecture

```
Sonor API (api.sonor.io)                   Signal API (signal.sonor.io)
├─ GET /api/public/llms/data ◄──────────── AI generates managed_llm_schema
├─ GET /api/public/llms/txt                + llms_public_summary per page
├─ GET /seo/llms/preview (auth'd)          via seo-meta-optimization pipeline
└─ GET /seo/llms/analytics (auth'd)
         │
         ▼
site-kit (npm: @sonordev/site-kit/llms)
├─ generateLLMsTxt()       ← builds markdown from Sonor data
├─ createLLMsTxtHandler()  ← zero-config Next.js route handler
├─ buildAiDiscoveryHeaders() ← Link header for crawler discovery
├─ createProxy({ llmsDiscovery }) ← sends the Link header from proxy.ts
├─ buildAiCrawlerRules() / createRobotsTxtHandler() ← robots.txt for AI crawlers
├─ writeLLMsTxtToPublic()  ← build-time static file generation
├─ createLlmsRevalidateHandler() ← on-demand ISR cache bust
├─ LLMSchema (RSC)         ← JSON-LD with managed_llm_schema + isPartOf
├─ SpeakableSchema         ← JSON-LD for voice assistants
└─ AEO* components         ← semantic HTML for AI extraction
         │
         ▼
Next.js Site
├─ /llms.txt         ← route handler (dynamic or static)
├─ /llms-full.txt    ← extended version (more pages/FAQ)
├─ Link header       ← rel="describedby" on all HTML responses
├─ robots.txt        ← allows AI crawlers to access /llms.txt
└─ JSON-LD in <head> ← managed_llm_schema per page
```

### Contract

The GEO system uses a shared contract (`@sonordev/site-kit/llms/contract`) consumed by site-kit, Sonor API, and Signal API. The contract defines:

- **`LLM_GEO_CONTRACT_VERSION`** (currently `1`) — increment on breaking payload changes
- **`LLMS_PUBLIC_SUMMARY_MAX_LENGTH`** (`400`) — max chars for page link notes
- **Sanitizers** — `sanitizeLlmsPublicSummary()`, `sanitizeLlmsDisclaimerLine()`, `sanitizePrimaryLanguageTag()`, `pickManagedLlmSchemaForJsonLd()`

See `LLM_GEO_CONTRACT.md` for the full spec.

***

## Implementation Guide

### Step 1: Route Handlers (required)

```ts
// app/llms.txt/route.ts
import { createLLMsTxtHandler } from '@sonordev/site-kit/llms'
export const GET = createLLMsTxtHandler()
// Next 15+: GET handlers are dynamic by default — opt into prerendering
export const revalidate = 3600
```

```ts
// app/llms-full.txt/route.ts
import { createLLMsFullTxtHandler } from '@sonordev/site-kit/llms'
export const GET = createLLMsFullTxtHandler()
export const revalidate = 3600
```

Both handlers:

- Serve static `public/llms.txt` if it exists (build-time optimized, `preferStatic: true` default)
- Fall back to dynamic generation from Sonor API
- Set `Cache-Control` with `s-maxage=3600` and `stale-while-revalidate=86400`
- Return a weak `ETag` (conditional `If-None-Match`/304 handling happens at the Next static layer / CDN)
- Include `X-Generated-At` and `X-Sections` response headers
- Never read the incoming `Request`, so the route prerenders statically (○) instead of
  bailing out with a "Dynamic server usage" error during `next build`

### Step 2: Discovery header (required)

**Proxy (recommended)**

```ts
// proxy.ts
import { createProxy } from '@sonordev/site-kit/proxy'

export default createProxy({
  llmsDiscovery: { siteUrl: 'https://example.com' },
})

// Inlined: an imported `config.matcher` is a build error. See the proxy README.
export const config = {
  matcher: [
    '/((?!_next/static|_next/image|favicon\\.ico|.*\\.(?:ico|png|jpg|jpeg|gif|webp|svg|woff2?)$).*)',
  ],
}
```

This adds `Link: <https://example.com/llms.txt>; rel="describedby"; type="text/markdown"`
to every request that could be for a page: GET or HEAD, with an Accept header
that allows HTML or no Accept header at all. Requests with no Accept header
count because that's how curl and many crawlers and AI fetchers ask, and
they're who the header is for (they used to be skipped). Skipped:
file-like paths (`/llms.txt`, `/sitemap.xml`), `/api/` routes, Accept headers
that name only non-HTML types, and Next's RSC navigation requests. The rule is
`wantsLlmsDiscoveryLink`, exported from `@sonordev/site-kit/llms`.

**No proxy? next.config**

```ts
// next.config.ts
import { withSiteKitConfig } from '@sonordev/site-kit/config'

export default withSiteKitConfig({ llmsTxtDiscoveryLink: true })
```

This sends the same header from next.config `headers()`, on every route. It
reads the origin from `NEXT_PUBLIC_SITE_URL`. If you also set
`nextConfig.headers`, yours replaces the helper's, so merge
`buildAiDiscoveryHeaders({ siteUrl })` into it yourself.

**Not from the root layout.** Older versions of this README showed
`export async function headers()` in `app/layout.tsx`. That isn't a Next API:
`headers()` is a next.config option, and exported from a layout it's an
ordinary function nobody calls, so it sends nothing. `sonor-setup geo` fails a
site wired that way and says so. Move it to the proxy.

`buildAiDiscoveryHeaders` builds the header value for either place, including
multilingual alternates:

```ts
buildAiDiscoveryHeaders({
  siteUrl: 'https://example.com',
  languageAlternates: [
    { hreflang: 'fr', href: 'https://example.com/fr/llms.txt' },
  ],
})
```

### Step 3: robots.txt (required)

Name the AI crawlers explicitly with the kit's curated lists rather than a
hand-rolled one, so a crawler the kit adds reaches every site on its next
upgrade:

```ts
// app/robots.ts
import type { MetadataRoute } from 'next'
import { buildAiCrawlerRules } from '@sonordev/site-kit/llms'

export default function robots(): MetadataRoute.Robots {
  return {
    rules: [
      { userAgent: '*', allow: '/', disallow: ['/api/'] },
      // Retrieval crawlers (they cite you) and training crawlers, each in its
      // own group. training: 'block' keeps retrieval and opts out of training.
      ...buildAiCrawlerRules({ disallow: ['/api/'] }),
    ],
    sitemap: 'https://example.com/sitemap.xml',
  }
}
```

A crawler that matches a named group ignores the `*` group, so pass the same
`disallow` to both. The lists are `AI_RETRIEVAL_CRAWLERS` and
`AI_TRAINING_CRAWLERS`. Don't keep a local list beside them; if an agent is
missing, add it to `src/llms/aiRobots.ts`.

Don't add a `Content-Signal` line (contentsignals.org). Google's robots.txt
parser reports it as an "Unknown directive" error in Search Console, and
robots.txt is the one file every crawler has to parse cleanly. Express the
same preference with `buildAiCrawlerRules({ training: 'allow' | 'block' })`.
`createRobotsTxtHandler` (a plain-text route handler) ignores
`contentSignals` since 6.3.4; remove it from a site's config when you next
touch the file.

These helpers are exported from `@sonordev/site-kit/llms`, their home.
`@sonordev/site-kit/robots` re-exports them too, beside `createRobots`: same
functions, either import works.

### Step 4: Keep llms.txt out of the sitemap and the index

An XML sitemap lists indexable HTML pages. llms.txt is a plain-text restatement
of those pages, so listing it invites search engines to index a thin duplicate
of the site. AI crawlers don't need it listed: they fetch `/llms.txt` by
convention, and the `Link: rel="describedby"` discovery header points at it from
every page. Leave `includeLlmsTxtInSitemap` / `includeLlmsFullTxtInSitemap` off
(the default); `sonor-setup geo` warns when a sitemap lists them.

`createLLMsTxtHandler` and `createLLMsFullTxtHandler` send
`X-Robots-Tag: noindex` by default, which keeps the file out of search indexes
without stopping AI crawlers from fetching it. Pass `noindex: false` only if you
want it in search results.

A static `public/llms.txt` (the build-time write below) is served straight from
the CDN, so neither the handler nor next.config `headers()` ever sees it. Set the
header in host config instead:

```toml
# netlify.toml
[[headers]]
  for = "/llms.txt"
  [headers.values]
    X-Robots-Tag = "noindex"

[[headers]]
  for = "/llms-full.txt"
  [headers.values]
    X-Robots-Tag = "noindex"
```

```ts
createSitemap({
  baseUrl: 'https://example.com',
  optimizedLLMsTxt: true,                // Write build-time static file (default: true)
})
```

The build-time write never persists a failure stub: if the Sonor fetch fails and no local data is available, `writeLLMsTxtToPublic()` skips the write (warning logged) so an existing `public/llms.txt` / `public/llms-full.txt` keeps serving via `preferStatic`.

Sonor generates this file with an LLM call that can take about a minute, and a
prerendered route only gets 60s, so the in-route write often times out and keeps
the previous file. For a refresh on every build, let the postbuild own it:

```jsonc
"scripts": { "postbuild": "sonor-register-sitemap --write-llms" }
```

with `optimizedLLMsTxt: false` in `createSitemap`, so one writer owns the file.
Reads behind this write always bypass Next's Data Cache — hosts persist it
between builds, and a cached read is how a site shipped a months-old llms.txt.

### Step 5: On-Demand Revalidation (optional)

When Sonor data changes (page summaries, FAQs, etc.), bust the ISR cache instantly instead of waiting for `s-maxage` to expire:

```ts
// app/api/revalidate-llms/route.ts
import { createLlmsRevalidateHandler } from '@sonordev/site-kit/llms'
export const POST = createLlmsRevalidateHandler(process.env.REVALIDATION_SECRET!)
```

Sonor can POST to this endpoint with `Authorization: Bearer <secret>` to revalidate `/llms.txt` and `/llms-full.txt`.

#### Sonor's SEO webhook (`/api/seo-revalidate`)

When a title, description or schema changes in Sonor, Sonor POSTs the
affected paths and cache tags to the site with the project key.
`createSeoRevalidationHandler` wraps the handler above and regenerates those
pages, `/sitemap.xml` and both llms files without a rebuild:

```ts
// app/api/seo-revalidate/route.ts
import { revalidatePath, revalidateTag } from 'next/cache'
import { createSeoRevalidationHandler } from '@sonordev/site-kit/llms'

export const runtime = 'nodejs'

export async function POST(request: Request) {
  return createSeoRevalidationHandler({
    secret: process.env.SONOR_API_KEY || '',
    revalidatePath,
    revalidateTag,
    publicationBasePath: '/insights', // only when the site has a publication
  })(request)
}
```

- Auth is `Authorization: Bearer <SONOR_API_KEY>` only, compared in constant
  time. No `?secret=` query form.
- Body: `{ paths?, path?, tags?, tag?, revalidateAll? }`, at most 16 KB and
  100 paths/tags. Every path must be a same-site local path (no scheme, `//`,
  query, fragment, `[segment]` or dot segment, even percent-encoded). One bad
  entry refuses the whole call (400) before any cache is touched.
- `revalidateAll`, or a tag-only call carrying `seo`, also revalidates the root
  layout. Tags expire immediately (`{ expire: 0 }`).
- `publicationBasePath` refreshes the publication index plus `rss.xml` and
  `feed.xml` on every call, and maps legacy `/blog/...` paths to the same URL
  under the publication root, refreshing both.
- `secret` can be a getter (`() => process.env.SONOR_API_KEY`), read on every
  call, so the handler can be created once at module scope.
- `extraPaths` are regenerated on every call. `extendPayload(payload, body)`
  adds paths or tags from body fields the handler doesn't read; its result is
  validated like the body. `@sonordev/agency-site-kit/revalidate` uses both
  for portfolio hubs, `slug`/`slugs` and its default `portfolio` tag.

***

## llms.txt Generation

### How `generateLLMsTxt()` Works

1. Fetches all data via `getLLMsData()` → `GET /api/public/llms/data`
2. Optionally merges with local data (`getLocalData` callback) when Sonor returns empty
3. Builds markdown sections in order: **Header → About → Services → Portfolio → Contact → FAQ → Pages → Optional → Full Context Link → Knowledge Graph → Topic Clusters → Custom Sections**
4. Portfolio, Entity, and Topic Cluster sections are fetched **in parallel** via `Promise.allSettled()`
5. Returns `{ markdown, metadata }` where metadata includes `sections`, `attempted_sections`, and `failed_sections`

### llms.txt Spec Compliance (llmstxt.org)

```markdown
# Business Name

> Tagline or summary
> Content index last updated: 2026-04-08T00:00:00Z
> Primary language: en
> This information is provided for reference purposes only.

## About

Business description...

## Services

- [Service Name](/services/slug): Brief description

## Portfolio & Case Studies

### [Project Title](https://example.com/work/project)

Description of the project.

## Contact Information

- **Phone:** 555-1234
- **Email:** info@example.com
- **Address:** 123 Main St, City, State

## Frequently Asked Questions

### How do you work?

We follow a proven process...

## Site Pages

- [Home](https://example.com/): Public-safe summary of the page content
- [About](https://example.com/about): Learn about our team and mission

## Optional

- [Privacy Policy](https://example.com/privacy): How we collect and use your information

## Full context

- [llms-full.txt](https://example.com/llms-full.txt): Expanded index for large-context systems.

## Knowledge Graph

### Business Name (Primary)

- **Type:** Organization
- **Schema:** LocalBusiness

## Topic Clusters

### Family Law Basics
Topic: Family Law
Area: Springfield
Articles: 8
Service page: https://example.com/services/family-law
Pillar: [Complete Guide to Family Law](/article/family-law-guide)
- [How to File for Divorce](/article/filing-divorce) (Mar 15, 2026)
- [Child Custody Laws](/article/custody-laws) (Mar 10, 2026)
```

### Configuration Options

```ts
generateLLMsTxt({
  // Section toggles (all default to true)
  includeBusinessInfo: true,
  includeServices: true,
  includeFAQ: true,
  includePages: true,
  includeContact: true,
  includePortfolio: true,
  includeEntities: true,

  // Limits
  maxFAQItems: 20,          // default 20 (100 in full mode)
  maxPages: 50,             // default 50 (200 in full mode)
  maxPortfolioItems: 20,    // default 20 (50 in full mode)
  maxEntities: 50,          // default 50 (200 in full mode)
  maxArticlesPerCluster: 5, // default 5

  // Spec features
  linkToFullLlms: true,                       // Append "Full context" section
  optionalPagePaths: ['/privacy', '/terms'],   // Moved from Site Pages to ## Optional
  pageListNotesFromPublicSummaryOnly: false,   // Strict mode: no description fallback

  // Overrides (take precedence over Sonor settings)
  headerPrimaryLanguage: 'en',
  headerDisclaimer: 'This information is for reference purposes only.',

  // Custom sections
  customSections: [{ title: 'Specializations', content: 'We specialize in...' }],

  // Local data fallback when Sonor is empty
  getLocalData: async () => ({ business: {...}, services: [...], ... }),
})
```

### Optional pages

`optionalPagePaths` demotes pages to `## Optional`, the section llmstxt.org lets a
short-context parser skip. A matching page **moves**: it's removed from
`## Site Pages` and listed once under Optional with the same URL and note. The
filter runs before `maxPages`, so demoting a page frees an index slot for the next
one. `/privacy`, `privacy` and `/privacy/` all match the same page. A path with no
matching page is still listed under Optional. The section needs a resolvable base
URL (see `baseUrl`); without one it's skipped and the pages stay in the index
rather than disappearing.

### Metadata Returned

```ts
const { markdown, metadata } = await generateLLMsTxt({})

metadata.generated_at      // ISO timestamp
metadata.project_id        // From API key
metadata.sections          // ['header', 'about', 'services', 'faq', 'pages', ...]
metadata.attempted_sections // ['portfolio', 'knowledge-graph', 'topic-clusters']
metadata.failed_sections   // ['knowledge-graph'] — if entity API was unavailable
```

***

## HTTP Caching

All llms.txt responses use this header strategy:

```
Content-Type: text/plain; charset=utf-8
Cache-Control: public, max-age=300, s-maxage=3600, stale-while-revalidate=86400, stale-if-error=86400
ETag: W/"<sha1-base64url>"
Vary: Accept-Encoding
```

- **Browser:** caches 5 minutes, then revalidates
- **CDN:** caches 1 hour, serves stale for 24 hours while revalidating
- **304 support:** The Next static layer / CDN matches `If-None-Match` against the emitted
  ETag. The handlers themselves never read request headers — doing so would force the
  route dynamic and break static prerendering

The `llmsResponseHeaders(body, extra?)` helper is exported for custom route handlers.

***

## JSON-LD Integration

### LLMSchema (via ManagedSchema)

The `LLMSchema` component in `@sonordev/site-kit/seo` emits `managed_llm_schema` from Sonor as JSON-LD:

```html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "WebPage",
  "name": "Divorce Law",
  "description": "Comprehensive divorce representation...",
  "url": "https://example.com/divorce",
  "additionalType": "https://sonor.io/ns/LLMOptimizedContent",
  "isPartOf": {
    "@type": "WebSite",
    "@id": "https://example.com/#website",
    "url": "https://example.com"
  }
}
</script>
```

- `managed_llm_schema` is generated by Signal AI during SEO meta optimization
- Only **known keys** from `MANAGED_LLM_SCHEMA_KNOWN_KEYS` are emitted (safety filter)
- `isPartOf` uses a stable `@id` pattern: `${siteUrl}/#website`

### Organization Stub

`createWebSiteOrganizationStub()` from `@sonordev/site-kit/seo` creates paired Organization + WebSite entities for `@graph` injection with stable `@id` anchors (`/#organization`, `/#website`).

***

## AEO Components

Semantic HTML components with schema.org microdata and `data-sonor-*` attributes for AI extraction.

| Component              | Schema Type     | Use Case                         |
| ---------------------- | --------------- | -------------------------------- |
| `AEOBlock`             | Question/Answer | FAQ-style Q\&A content           |
| `AEOSummary`           | —               | Key points lists (speakable)     |
| `AEODefinition`        | DefinedTerm     | Term/definition pairs            |
| `AEOSteps` / `AEOStep` | HowTo           | Step-by-step processes           |
| `AEOComparison`        | —               | Feature/option comparison tables |
| `AEOClaim`             | Claim           | Source-attributed factual claims |
| `AEOEntity`            | —               | Inline entity annotations        |
| `AEOProvenanceList`    | —               | Citation source lists            |
| `AEOCitedContent`      | —               | Content with numbered citations  |

All components support:

- `speakable` prop → adds `data-speakable="true"` for voice assistants
- `entityId` prop → links to knowledge graph via `data-sonor-entity`
- `className` prop → custom styling

### Example: Service Page with AEO

```tsx
import { AEOSummary, AEOSteps, AEOStep, AEOBlock } from '@sonordev/site-kit/llms'

export default function DivorcePage() {
  return (
    <article>
      <h1>Divorce Law Services</h1>

      <AEOSummary
        title="Key Facts"
        points={[
          '25+ years of family law experience',
          'Serving Springfield and the surrounding counties',
          'Free initial consultation available',
        ]}
        speakable
      />

      <AEOSteps title="The Divorce Process" speakable>
        <AEOStep name="Consultation" text="Meet with attorney to discuss your situation" position={1} />
        <AEOStep name="File Petition" text="Submit divorce petition to circuit court" position={2} />
        <AEOStep name="Negotiate" text="Work toward fair settlement" position={3} />
        <AEOStep name="Finalize" text="Court issues final decree" position={4} />
      </AEOSteps>

      <AEOBlock type="answer" question="How long does a divorce take?" speakable>
        Uncontested divorces typically take 60-90 days. Contested cases take 6-12 months.
      </AEOBlock>
    </article>
  )
}
```

***

## Speakable Schema

Marks page sections for voice assistant extraction via `SpeakableSpecification` JSON-LD.

```tsx
import { SpeakableSchema } from '@sonordev/site-kit/llms'

<SpeakableSchema
  type="WebPage"
  name="Divorce Law"
  url="https://example.com/divorce"
  speakable={{ cssSelectors: ['h1', '[data-speakable="summary"]'] }}
/>
```

Default selectors by page type:

| Type      | Selectors                                                                 |
| --------- | ------------------------------------------------------------------------- |
| `page`    | `h1`, `[data-speakable]`, `.intro`, `[role="main"] > p:first-of-type`     |
| `article` | `h1`, `.article-summary`, `article > p:first-of-type`, `[data-speakable]` |
| `service` | `h1`, `.service-description`, `[data-speakable="summary"]`                |
| `faq`     | `.faq-question`, `[data-speakable]`                                       |
| `contact` | `h1`, `.contact-info`, `[itemprop="address"]`                             |

***

## Sonor API Endpoints

### Public (API key auth via `x-api-key`)

| Endpoint                                          | Purpose                                                                                                                      |
| ------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| `GET /api/public/llms/data`                       | All LLM visibility data (business, services, FAQ, pages, meta)                                                               |
| `GET /api/public/llms/txt`                        | AI-optimized llms.txt markdown (calls Signal AEO, with fallback)                                                             |
| `GET /api/public/llms/txt?full=true`              | Extended version (200 pages, 100 FAQ)                                                                                        |
| `GET /api/public/llms/txt?publicSummaryOnly=true` | Strict mode: page notes from `llms_public_summary` only                                                                      |
| `GET /api/public/llms/business`                   | Business info only                                                                                                           |
| `GET /api/public/llms/services`                   | Services list                                                                                                                |
| `GET /api/public/llms/faq`                        | FAQ items (`getFAQItems(projectId?, limit?, site?)` sends `?site=`, so a microsite gets its own FAQs plus project-wide ones) |
| `GET /api/public/llms/pages`                      | Page summaries                                                                                                               |

### Authenticated (dashboard)

| Endpoint                                        | Purpose                           |
| ----------------------------------------------- | --------------------------------- |
| `GET /seo/projects/:id/llms/preview`            | Preview llms.txt markdown + stats |
| `GET /seo/projects/:id/llms/analytics?period=7` | AI crawler request breakdown      |

### Key Database Fields

| Table               | Column                    | Purpose                                                     |
| ------------------- | ------------------------- | ----------------------------------------------------------- |
| `seo_pages`         | `llms_public_summary`     | Public-safe summary for llms.txt link notes (max 400 chars) |
| `seo_pages`         | `managed_llm_schema`      | JSON-LD object for per-page LLM optimization                |
| `seo_pages`         | `language_alternates`     | Optional hreflang map (locale → URL)                        |
| `seo_pages`         | `llm_schema_generated_at` | When Signal last generated the schema                       |
| `projects.settings` | `primary_language`        | BCP 47 tag for llms.txt blockquote                          |
| `projects.settings` | `llms_disclaimer`         | Optional disclaimer line in blockquote                      |
| `llms_request_log`  | `bot_class`               | AI crawler classification for analytics                     |

***

## Environment Variables

```bash
# Required (server-only — SiteKitLayout injects into client automatically):
SONOR_API_KEY=sonor_xxxxxxxx_xxxxx

# Optional:
SONOR_API_URL=https://api.sonor.io          # Default
NEXT_PUBLIC_SITE_URL=https://example.com     # For CLI status checks
REVALIDATION_SECRET=your_secret              # For on-demand revalidation endpoint
```

***

## CLI Validation

`npx sonor-setup status` runs a comprehensive llms.txt health check when `NEXT_PUBLIC_SITE_URL` is set:

- **HTTP status** — must be 200
- **Content structure** — H1 title, blockquote summary, H2 section count
- **Response headers** — Content-Type, Cache-Control (s-maxage), ETag presence
- **Freshness** — parses `last_updated` from blockquote, warns if > 7 days old

***

## Exports

```ts
// Types
import type {
  LLMBusinessInfo, LLMContactInfo, LLMService, LLMFAQItem,
  LLMPageSummary, LLMPortfolioItem, LLMsDataResponse, LLMsPayloadMeta,
  GenerateLLMSTxtOptions, LLMSTxtContent, WriteLLMsTxtOptions,
  SpeakableConfig, SpeakableSchemaProps,
  AEOBlockProps, AEOSummaryProps, AEODefinitionProps,
  AEOClaimProps, AEOEntityProps, ContentProvenance,
  AEOProvenanceListProps, AEOCitedContentProps,
  AiDiscoveryHeadersOptions, LlmsLanguageAlternate,
} from '@sonordev/site-kit/llms'

// Contract (lightweight — safe for API imports)
import {
  LLM_GEO_CONTRACT_VERSION, LLMS_PUBLIC_SUMMARY_MAX_LENGTH,
  LLMS_DISCLAIMER_MAX_LENGTH, MANAGED_LLM_SCHEMA_KNOWN_KEYS,
  sanitizeLlmsPublicSummary, sanitizeLlmsDisclaimerLine,
  sanitizePrimaryLanguageTag, pickManagedLlmSchemaForJsonLd,
} from '@sonordev/site-kit/llms/contract'

// Generation
import { generateLLMsTxt, generateLLMsFullTxt } from '@sonordev/site-kit/llms'

// Route handlers
import {
  createLLMsTxtHandler, createLLMsFullTxtHandler, llmsResponseHeaders,
} from '@sonordev/site-kit/llms'

// Discovery
import { buildAiDiscoveryHeaders } from '@sonordev/site-kit/llms'

// API data fetchers (React cache()-wrapped)
import {
  getLLMsData, getBusinessInfo, getServices, getFAQItems,
  getPageSummaries, getOptimizedLLMsTxt,
} from '@sonordev/site-kit/llms'

// Build-time
import { writeLLMsTxtToPublic } from '@sonordev/site-kit/llms'

// Revalidation
import { createLlmsRevalidateHandler } from '@sonordev/site-kit/llms'

// Speakable
import {
  SpeakableSchema, createSpeakableSchema,
  getSpeakableSelectorsForPage, DEFAULT_SPEAKABLE_SELECTORS,
} from '@sonordev/site-kit/llms'

// AEO Components
import {
  AEOBlock, AEOSummary, AEODefinition, AEOSteps, AEOStep,
  AEOComparison, AEOClaim, AEOEntity, AEOProvenanceList, AEOCitedContent,
} from '@sonordev/site-kit/llms'

// Proxy (separate import path)
import { createProxy } from '@sonordev/site-kit/proxy'
// Config type: { llmsDiscovery?: { siteUrl: string; llmsPath?: string } | false }
```

***

## Testing

```bash
# Run all llms tests (42 tests across 4 suites)
pnpm test src/llms/

# Test suites:
# - contract.test.ts       — sanitizers, PII filtering, schema key allowlist
# - generateLLMsTxt.test.ts — generation: sections, limits, flags, fallbacks, metadata
# - handlers.test.ts       — response headers, ETag, Cache-Control
# - discovery-headers.test.ts — Link header format, hreflang validation
```

### Manual verification

```bash
# Check llms.txt
curl -s https://example.com/llms.txt | head -20

# Check headers
curl -sI https://example.com/llms.txt | grep -E 'ETag|Cache-Control|Content-Type|Link'

# Check discovery Link header on HTML page
curl -sI -H 'Accept: text/html' https://example.com/ | grep Link

# Check 304 support (served by the CDN/static layer, not the handler)
ETAG=$(curl -sI https://example.com/llms.txt | grep ETag | awk '{print $2}' | tr -d '\r')
curl -sI -H "If-None-Match: $ETAG" https://example.com/llms.txt | head -1
# Should return: HTTP/2 304
```
