diff --git a/docs/roadmap/cache.md b/docs/roadmap/cache.md deleted file mode 100644 index ce7a0dc..0000000 --- a/docs/roadmap/cache.md +++ /dev/null @@ -1,135 +0,0 @@ -# Prompt Cache Control Roadmap - -## Purpose - -Scriptorium should support provider prompt-cache controls for OpenAI-compatible gateways that expose Anthropic-style cache breakpoints, especially OpenRouter. - -The feature should preserve existing prompt definitions. Prompt authors should opt in with optional message-level cache metadata, and Scriptorium should report cache usage when compatible providers return it. - -Implementation steps belong in `docs/roadmap/implementation.md`. - -## Target State - -Prompt authors can mark a rendered message as an explicit cache breakpoint: - -```yaml -messages: - - role: system - content_file: ../common/transcript.system.md - cache_control: - type: ephemeral - ttl: 1h -``` - -Existing messages without `cache_control` continue to render and serialize as string content: - -```json -{ - "role": "system", - "content": "rendered text" -} -``` - -Messages with `cache_control` serialize as a single text content block: - -```json -{ - "role": "system", - "content": [ - { - "type": "text", - "text": "rendered text", - "cache_control": { - "type": "ephemeral", - "ttl": "1h" - } - } - ] -} -``` - -Scriptorium should parse and expose provider cache usage when compatible response fields are present, including cached prompt tokens and cache-write tokens. - -## Prompt Authoring Policy - -Cache breakpoints should be used for stable reusable prompt prefixes. - -Recommended ordering: - -1. Put stable, reusable context first. -2. Put `cache_control` on the last stable message that should be part of the reusable prefix. -3. Put per-run dynamic inputs after that breakpoint. - -Example: - -```yaml -messages: - - role: system - content_file: ../common/transcript.system.md - - role: user - content_file: ./character_meta_analysis.task.md - - role: user - content_file: ./character_meta_analysis.instructions.md - cache_control: - type: ephemeral - ttl: 1h - - role: user - content: | - <<>> - - role: user - content: | - <<>> -``` - -## Supported Cache-Control Shape - -Initial support is message-level only: - -```yaml -cache_control: - type: ephemeral - ttl: 1h -``` - -Rules: - -- `cache_control` is optional on each message. -- `cache_control.type` is required when `cache_control` is present. -- The only supported `type` value is `ephemeral`. -- `ttl` is optional. -- When set, the only supported `ttl` value is `1h`. -- Empty `ttl` is omitted from the outbound payload. -- Prompt decoding remains strict; unknown cache-control fields are rejected. -- Existing `content` and `content_file` rules remain unchanged. - -## Compatibility - -This feature should be backward compatible for existing prompt definitions. - -Compatibility requirements: - -- Existing prompt YAML without `cache_control` loads unchanged. -- Existing `render` output remains valid. -- Existing outbound request payloads remain string-content messages unless `cache_control` is configured. -- Existing integrations do not need to send new request fields. -- CLI and HTTP callers do not need new request options for the initial feature. - -The only intentional prompt-definition contract change is the new optional `messages[].cache_control` object. - -## Deferred Work - -These are intentionally out of scope for the initial feature: - -- `content_blocks` prompt syntax. -- Multiple text blocks inside a single message. -- Image, tool, or non-text content blocks. -- Provider-specific automatic prompt caching toggles. -- Top-level OpenRouter `cache_control`. -- General-purpose serialization of `extra_params`. -- Provider-specific validation profiles for cache-control limits. - -These can be added later without changing the message-level cache-control contract. diff --git a/docs/roadmap/implementation.md b/docs/roadmap/implementation.md deleted file mode 100644 index 3131d0a..0000000 --- a/docs/roadmap/implementation.md +++ /dev/null @@ -1,390 +0,0 @@ -# Prompt Cache Control Implementation Plan - -## Scope - -This plan implements the feature defined in `docs/roadmap/cache.md`. - -The implementation target is message-level Anthropic/OpenRouter-style cache control for rendered prompt messages, plus cache usage observability. Do not implement deferred roadmap items such as `content_blocks`, top-level `cache_control`, or general `extra_params` serialization. - -Follow the policy documents under `docs/policy/` while implementing: - -- Keep prompt/profile/HTTP/config decoding strict. -- Keep adapters thin. -- Keep provider request serialization in `internal/llm`. -- Keep prompt loading in `internal/promptdef`. -- Update canonical non-roadmap docs only when the behavior is implemented. -- Do not accept or emit raw secret values. - -## Stage 1: Domain Types And Prompt Loader - -Objective: make prompt YAML accept and validate optional `messages[].cache_control`, and carry normalized cache-control metadata in domain prompt definitions. - -Files to update: - -- `internal/domain/domain.go` -- `internal/promptdef/filesystem_repository.go` -- `internal/promptdef/repository_test.go` -- `internal/promptdef/testdata/` - -Domain decisions: - -- Add a typed cache-control value in `internal/domain`. -- Use a pointer on message structs so absence is distinguishable from an empty object. -- Use `omitempty` on JSON tags for the new field so existing no-cache prompt and render hashes do not gain `null` fields. - -Use this domain shape: - -```go -type CacheControlType string - -const ( - CacheControlEphemeral CacheControlType = "ephemeral" -) - -type CacheControl struct { - Type CacheControlType `yaml:"type" json:"type"` - TTL string `yaml:"ttl,omitempty" json:"ttl,omitempty"` -} -``` - -Add cache control to: - -```go -type PromptMessageTemplate struct { - Role string `yaml:"role"` - Content string `yaml:"content"` - ContentFile string `yaml:"content_file"` - CacheControl *CacheControl `yaml:"cache_control,omitempty" json:"cache_control,omitempty"` -} -``` - -Prompt loader changes: - -- Add this field to `promptMessageFile`: - -```go -CacheControl *cacheControlFile `yaml:"cache_control"` -``` - -- Add a private `cacheControlFile` struct with `type` and `ttl`. -- In `normalizePromptDefinition`, validate `cache_control` per message. -- Trim whitespace from `type` and `ttl`. -- Reject present `cache_control` when `type` is empty. -- Reject any `type` other than `ephemeral`. -- Accept omitted or empty `ttl`. -- Reject any non-empty `ttl` other than `1h`. -- Copy normalized values into `domain.PromptMessageTemplate.CacheControl`. -- Preserve the existing rule that each message must set exactly one of `content` or `content_file`. - -Testing requirements: - -- Add a valid fixture with `cache_control: {type: ephemeral, ttl: 1h}`. -- Add a valid fixture with `cache_control: {type: ephemeral}`. -- Add invalid fixtures for: - - empty `cache_control.type` - - unsupported `cache_control.type` - - unsupported `cache_control.ttl` - - unknown field under `cache_control` -- Assert valid fixtures load with the expected `CacheControl` values. -- Assert invalid fixtures fail with `promptdef.ErrInvalidPromptDefinition` or `promptdef.ErrInvalidYAML` as appropriate. -- Re-run: - -```bash -go test ./internal/promptdef -``` - -Do not update renderer, formatter, LLM serialization, adapters, or user-facing docs in this stage. - -## Stage 2: Rendering, Prepared Output, And Render Hashing - -Objective: preserve cache-control metadata through rendering, expose it in prepared-run output, and make `rendered_prompt_hash` reflect outbound-affecting cache metadata. - -Files to update: - -- `internal/domain/domain.go` -- `internal/prompt/go_renderer.go` -- `internal/prompt/renderer_test.go` -- `internal/usecase/runner.go` -- `internal/usecase/runner_test.go` -- `internal/domain/prepared_run_test.go` -- `internal/format/prepared_run.go` -- `internal/format/prepared_run_test.go` - -Domain changes: - -- Add cache control to rendered messages: - -```go -type RenderedMessage struct { - Role string `json:"role"` - Content string `json:"content"` - CacheControl *CacheControl `json:"cache_control,omitempty"` -} -``` - -Renderer changes: - -- When a `PromptMessageTemplate` has `CacheControl`, copy it to the returned `RenderedMessage`. -- Do not mutate the template's pointer in place. -- Add and use a small helper such as `cloneCacheControl` to avoid aliasing. -- Do not change template execution behavior for `{{input "name"}}` or `{{.var}}`. - -Rendered hash decision: - -- `rendered_prompt_hash` must change when cache-control metadata changes, because the outbound provider request changes. -- Preserve existing hashes for no-cache prompts by keeping the current role/content hash flow. -- Append normalized cache-control metadata only when present, for example: - - `cache_control.type=` - - `cache_control.ttl=` only when non-empty -- Include cache-control fields in a stable order. - -Prepared-run formatting: - -- JSON prepared-run output should include `cache_control` on messages only when present. -- Text prepared-run output should show cache-control metadata under the specific message, before `content`, for example: - -```text - cache_control: ephemeral ttl=1h - content: | -``` - -- If `ttl` is empty, omit the `ttl=...` suffix. -- Do not print cache-control lines for messages without cache control. - -Testing requirements: - -- Renderer test: cache control is copied to rendered messages. -- Renderer test: changing returned rendered cache control does not mutate the source template. -- Runner test: `rendered_prompt_hash` changes when cache control differs and remains stable for identical rendered messages. -- Format JSON test: `cache_control` appears only when present. -- Format text test: cache control appears under the correct message. -- Existing render/format tests still pass. -- Re-run: - -```bash -go test ./internal/prompt ./internal/usecase ./internal/domain ./internal/format -``` - -Do not update outbound LLM serialization or adapter response surfaces in this stage. - -## Stage 3: Outbound Chat Serialization And Usage Parsing - -Objective: serialize cache-controlled messages as one-item text content-block arrays and parse cache usage fields from compatible provider responses. - -Files to update: - -- `internal/llm/openai_compatible_client.go` -- `internal/llm/openai_compatible_client_test.go` -- `internal/domain/domain.go` - -Request serialization decisions: - -- Keep string content for messages without cache control. -- Use content blocks only when `RenderedMessage.CacheControl` is non-nil. -- Do not introduce `content_blocks` prompt syntax. -- Do not serialize top-level `cache_control` or `extra_params`. -- Split request and response wire message structs. The current `openAIChatMessage` is used in both request and response paths; changing request content to `any` should not make response decoding less strict than necessary. - -Use request wire types equivalent to: - -```go -type openAIChatRequestMessage struct { - Role string `json:"role"` - Content any `json:"content"` -} - -type openAIChatTextContentBlock struct { - Type string `json:"type"` - Text string `json:"text"` - CacheControl *openAICacheControl `json:"cache_control,omitempty"` -} - -type openAICacheControl struct { - Type string `json:"type"` - TTL string `json:"ttl,omitempty"` -} -``` - -When cache control is present, content should serialize as: - -```json -[ - { - "type": "text", - "text": "rendered text", - "cache_control": { - "type": "ephemeral", - "ttl": "1h" - } - } -] -``` - -When `ttl` is empty, omit `ttl`. - -Usage parsing decisions: - -- Extend `domain.TokenUsage`: - -```go -type TokenUsage struct { - PromptTokens int - CompletionTokens int - TotalTokens int - CachedTokens int - CacheWriteTokens int -} -``` - -- Parse `usage.prompt_tokens_details.cached_tokens` into `CachedTokens`. -- Parse `usage.cache_write_tokens` into `CacheWriteTokens`. -- Treat absent cache fields as zero. -- Do not fail response decoding if cache fields are absent. -- If a provider later returns a different cache-write field, add it in a later change after observing a real response. - -Testing requirements: - -- Existing request serialization test still sees string message content for no-cache messages. -- New request serialization test sees content-block array for cache-controlled messages. -- New request serialization test verifies empty `ttl` is omitted. -- Response parsing test verifies cached and cache-write tokens are mapped. -- Response parsing test verifies absent cache usage fields remain zero. -- Existing structured-output and service-tier behavior remains unchanged. -- Re-run: - -```bash -go test ./internal/llm -``` - -Do not update CLI or HTTP observability in this stage unless required by compilation after extending `TokenUsage`. - -## Stage 4: CLI And HTTP Cache Usage Observability - -Objective: surface parsed cache usage through existing CLI and HTTP result paths without adding request knobs. - -Files to update: - -- `internal/adapter/cli/run.go` -- `internal/adapter/cli/run_test.go` -- `internal/adapter/http/dto.go` -- `internal/adapter/http/handler.go` -- `internal/adapter/http/handler_test.go` - -CLI decisions: - -- Do not add CLI flags. -- Keep success summaries concise. -- Preserve the existing summary output when cache usage is zero. -- When either cache field is non-zero, append: - -```text -cached_tokens= cache_write_tokens= -``` - -HTTP decisions: - -- Do not add request fields. -- Add numeric fields to `metadata.usage`: - - `cached_tokens` - - `cache_write_tokens` -- Always include these fields as numbers, matching the current non-omitempty numeric usage fields. -- Continue omitting raw model output unless `include_raw_output` is true. - -Testing requirements: - -- CLI summary omits cache fields when both are zero. -- CLI summary includes cache fields when either is non-zero. -- HTTP success response includes `metadata.usage.cached_tokens`. -- HTTP success response includes `metadata.usage.cache_write_tokens`. -- Existing HTTP request mapping is unchanged. -- Re-run: - -```bash -go test ./internal/adapter/cli ./internal/adapter/http -``` - -## Stage 5: Canonical Documentation - -Objective: update implemented-behavior docs after the feature works and tests pass. - -Files to update: - -- `docs/config.md` -- `docs/integrations/openai-compatible-chat.md` -- `docs/cli.md` -- `docs/troubleshooting.md` -- `docs/internal/runner.md` -- `docs/internal/adapters.md`, if adapter response behavior needs an explicit note -- `docs/roadmap/cache.md`, only if the accepted target changes during implementation - -Documentation requirements: - -- Keep all unimplemented ideas under `docs/roadmap/`. -- Document `messages[].cache_control` in the prompt definition reference in `docs/config.md`. -- Document the outbound message serialization rule in `docs/integrations/openai-compatible-chat.md`. -- Document parsed cache usage fields in `docs/integrations/openai-compatible-chat.md`. -- In `docs/cli.md`, note that prompt cache control is configured in prompt YAML, not with CLI flags. -- In `docs/troubleshooting.md`, add cache miss guidance: - - stable context should appear before the cache breakpoint - - dynamic input before the breakpoint changes the cache key - - providers may impose minimum token thresholds and breakpoint limits - - use cached/cache-write token counts to verify behavior -- Do not claim support for deferred roadmap items. - -Testing and verification: - -```bash -go test ./... -``` - -If documentation examples include new YAML snippets, ensure they match the strict prompt loader schema. - -## Stage 6: Final Verification - -Objective: verify the whole change set end to end. - -Required commands: - -```bash -go test ./... -go build ./cmd/scriptorium -``` - -Recommended smoke checks: - -```bash -go run ./cmd/scriptorium render \ - --config ./examples/config.yml \ - --prompt generic.markdown_summary \ - --input transcript=./examples/fixtures/transcript.md \ - --input glossary=./examples/fixtures/glossary.yml \ - --format json -``` - -If adding a new example prompt with `cache_control`, also smoke-render that prompt with `--format json` and inspect that the rendered message includes `cache_control`. - -Final review checklist: - -- Existing prompt definitions without `cache_control` still load. -- Prompt definitions with `cache_control` render successfully. -- Unknown cache-control YAML fields are rejected. -- No-cache outbound messages still serialize with string `content`. -- Cache-controlled outbound messages serialize with text content-block arrays. -- `rendered_prompt_hash` changes when cache-control metadata changes. -- Cache usage fields are zero when providers omit them. -- CLI and HTTP request surfaces have no new knobs. -- CLI and HTTP response surfaces expose cache usage. -- Non-roadmap docs describe only implemented behavior. -- `docs/roadmap/cache.md` still defines target state and deferred work, not task steps. - -## Stage Boundaries - -Stages 1 through 4 can be implemented in separate prompts if needed. Stage 5 should happen only after behavior is implemented. Stage 6 should be performed after all implementation and documentation changes are complete. - -If a stage fails because prior stage changes were not implemented, stop and complete the earlier stage first rather than papering over missing domain fields or DTO mappings. - -## Open Questions - -None required for the initial implementation. - -The plan chooses message-level `cache_control`, string content for legacy messages, text-block array content only when cache control is present, and no new CLI/HTTP request knobs. These choices match `docs/roadmap/cache.md` and preserve existing prompt-definition compatibility.