Keywords: autonomous search engine optimization; strategy compilation; nightly review; closed amendment vocabulary; context projection; operator configuration; longitudinal case study.
Scope and Claims
Our earlier publication described an architecture in which search, analytics, rank, local, technical and site evidence compile into a versioned strategy that constrains specialized agents.[2] It ended with an evaluation protocol and a boundary: no controlled study had yet measured the architecture, and the founding growth observation predated it.[1] This paper is the first longitudinal record against that protocol. It covers the ninety days from June 30 to September 27, 2026, on the eight client properties that ran the system in Full Autopilot, the mode in which the platform authors and activates strategy versions without a person approving each one.
The claims are bounded by design. We show that a strategy can be recompiled every night from every source the platform holds, that its amendments are bounded and auditable, that the same version reaches topic selection, the writer’s brief, the page and technical queues and the client portal, and that the operator’s own configuration now enters the review as evidence rather than as an invisible constraint. We show the search visibility of the eight properties over the ninety-day window, with each 28-day comparison window named. We do not claim that the architecture caused the visibility: a causal claim requires a declared treatment date and a prespecified counterfactual, and the design that supplies both is stated in the closing sections.[4]
Figure 1 shows the combined Search Console impressions of the eight properties over the ninety-day window, week by week, as the daily capture recorded them: 337,897 impressions and 1,302 clicks between June 30 and September 27, 2026. The roster and instrument facts that qualify the curve are stated in the caption and in Materials and Method.[9]
Summary of Contributions
The contributions of the period are listed in the order in which the paper documents them. Each is traceable to a versioned change, an executed contract test or a ledger read cited in the sources.[9][10][11][12]
- A longitudinal record against the evaluation protocol of the compiler paper: eight properties over ninety days, with every window and instrument named.
- Working rules for the nightly review, namely rotation by last review, one review per evidence state, and the rule that a version is a change, which ended the re-review loop in which sixty of sixty reviews had produced an activation.
- Strategy-aware topic selection. All three topic lanes read the active clusters and the open article backlog, the strategy’s own page-keyword rows enter the candidate list, and an accepted search phrase is attributed to the cluster whose query it matches.
- Operator configuration as review evidence, with a bounded path by which the review may ask the operator to examine a setting and no path by which it may rewrite one.
- Rank measured at the market city. Every reader of rank admits measured rows at the property’s rank scope, and the writer’s brief names the lane a number came from.
- A quarterly research clock anchored on the newest completed research run rather than on an activation timestamp that every nightly version re-stamped.
- A verification discipline, adversarial review of every change with execution of the real modules, which identified the eight defects documented below and pinned each correction with an executed test.
- A recomputable data module from which the search, rank and activity series plotted and tabulated in this paper are generated.
Background: From a Content Pipeline to a Living Strategy
The founding case study documented a feed-forward content pipeline: collect signals, choose a topic, draft, publish, repeat.[1] Its founding property grew more than twelvefold in impressions and its local counterpart reached a plateau within six weeks, and the note argued that the plateau, not the curve, was the more instructive result. Both properties ran the same kind of loop: the model chose a topic from cached evidence, wrote against a brief, and the outcome was measured afterward. Nothing in the loop decided what the site should be about next quarter, which pages should own which search intents, or when a piece of evidence was too old to act on.
The strategy compiler introduced that missing layer as a versioned authority.[2] The system described here is that architecture in operation, with one addition that turned a compiled artifact into a living one: a nightly review that reads the strategy against fresh evidence and proposes amendments in a closed vocabulary, and, under Full Autopilot, authors and activates the next version itself. The adjective is used in its literal sense: across the eight properties, 98 versions were activated in the ninety days ending September 30, 2026, and 28 of them were authored by the review rather than by research or a person.[9]
- 01
Observe
Search Console queries and pages, analytics sessions, paid rank and keyword data, competitor pages, local grid scans, crawl and index state, page speed, and the live site inventory.
- 02
Qualify
Each observation keeps its window, scope, method and source; a measured rank outranks a cached echo on the same day; a missing source stays absent rather than zero.
- 03
Compile
The active strategy binds clusters, owning pages, priorities and a typed backlog under one content hash; activation projects page ownership and tracked terms into operational tables.
- 04
Review nightly
A desktop turn assembles the strategy, the evidence blocks, the live state and, since September, the operator context; it proposes bounded amendments in a closed vocabulary.
- 05
Author and activate
Under Full Autopilot, accepted amendments become the next version only when they change something; identical proposals resolve without a version.
- 06
Project
Topic selection, the writer’s brief, page and technical queues, and the client portal read the same active version; each lane receives its own slice.
- 07
Measure
Published work, rank checks, search windows and the impact journal return as evidence for the next review under their own clocks.
Three differences from the content-only pipeline matter for what follows. First, the agents no longer carry the strategy in a prompt; they read a projection of the active version, so two agents working the same property cannot act on different versions of the strategy. Second, adaptation is a version, with a diff, a provenance stamp and an input digest, not an edit to an instruction. Third, the review sees evidence under the same eligibility rules the reports use: a measured rank outranks a cached echo on the same day, a missing source stays absent, and a comparison names its window.[13]
System Description: Evidence Sources and the Decision Path
The platform holds seven families of evidence for a property. Search Console contributes queries, pages, impressions, clicks and position by window. Analytics contributes sessions and organic landing pages. A paid data provider contributes keyword volume and difficulty, the site’s ranked keywords, result-page composition and competitor pages, and the weekly measured rank of a tracked keyword set at the property’s rank scope. Geographic grid scans contribute local map visibility from explicit coordinates. Audits contribute crawl, index, page-speed and schema state. A site inventory contributes the live pages, their canonical paths and internal structure. Research runs contribute a keyword universe, entities and competitor gaps. Each family updates on its own clock.[2]
These sources reach a decision through two compilers. The research pipeline compiles a strategy draft in named stages and validates its synthesis against evidence identifiers. The nightly review compiles a bounded input: nine data blocks serialized deterministically, hashed into an input digest, followed by two prose blocks that carry facts only. The reviewer, a language model running as a desktop turn at no marginal cost, returns findings typed as support, contradiction, gap or amendment, and an amendment may only add or remove a cluster query, set a charter or cluster field, retire a cluster, or add, re-prioritize or drop a backlog item. Anything outside that vocabulary is discarded before it can be applied.[10]
strategy
The active version, its backlog (with any automatic priority cap) and the last diff.
gsc_clusters
Each cluster’s own queries plus the strongest fresh queries no cluster holds.
rank
The measured rank panel at the project’s rank scope, with the cache echo kept apart.
competitor_gap
The research run’s competitor artifact, including a run cancelled at activation.
pagespeed · crawl_index
Technical state from the latest audits and index inspections.
impact_journal · monthly_snapshot
Measured before-and-after windows for shipped changes and the monthly figures.
algorithm_updates
Dated search-engine updates inside the comparison window.
operator_context
Seed terms with the cluster that holds each, banned topics, writer guidelines and the client summary, as facts and days.
live_state
The platform’s own derived reading of the strategy against the evidence it already holds.
The design follows a principle we have found more durable than any particular prompt: accurate context outperforms restrictions. Long-context studies show that a model does not reliably use the right fact when it sits in the middle of a long input, and our own earlier incident showed a correct file left unread because an index omitted it.[5][14] The review therefore receives a small, complete dossier with stable identifiers rather than the whole client state, and it is asked for decisions inside a vocabulary the applier can validate. Workflows carry the predictable steps; the model is asked only where semantic judgment is necessary.[6]
The projections are equally bounded. Topic selection sees the active clusters, up to six queries each with the owning page, and the open article backlog. The writer’s brief carries the accepted search phrase, the cluster it serves, up to five stored queries per cluster, the site’s standing on that phrase with its instrument named, and the angles recent posts already used. The client portal derives its next and shipped lists from client-facing channels only. The strategy is a single authority with several projections, and no projection is the whole.
Algorithm and Functionality Advances in the Period
The period’s algorithm and functionality advances were shipped as separate, reviewed versions between September 25 and September 30, 2026, each measured against the autoblog doctrine that governs the content lane: context and wording only, no new validator, gate, hold, skip, retry or tool restriction, and nothing that makes a post less likely to publish.[12] The advances are summarized in Table 1; the sections that follow document their effects.
| Date | Change | What it does |
|---|---|---|
| Sep 25 | Automatic authored changes | The nightly review may author and activate the next strategy version under Full Autopilot; a human draft is never overwritten; an owner’s explicit top priority is never claimed by the automation. |
| Sep 29–30 | Rank scope | The weekly rank panel measures a local business at its market city rather than at the country level; every reader of rank admits measured rows at that scope and keeps the daily cache echo apart. |
| Sep 30 | Working rules of the review | Projects rotate by last review; a review never re-reviews the version it just authored; amendments that already hold resolve their findings without a new version; each cluster’s own queries and the strongest unclustered queries appear in the input. |
| Sep 30 | The brief is built on the search phrase | The first dispatch of a queued headline carries the phrase the topic was accepted for, and the writer is told why that phrase belongs in the title, heading, slug and description. |
| Sep 30 | Nightly cap of eight | Every authoring project takes its turn each night instead of one in three; the turns are desktop time, not paid calls. |
| Sep 30 | The quarterly research clock | Deep research comes due ninety days after the newest completed run, not after an activation timestamp that every nightly version re-stamped. |
| Sep 30 | Owner-written summaries kept | A report attach never overwrites a client summary the owner wrote; the portal’s next and shipped lists carry client-facing channels only, in a deterministic order. |
| Sep 30 | Topic selection reads the strategy | All three topic lanes see the active clusters and the open article backlog; the strategy’s own page-keyword rows enter the candidate list through the existing fences; an accepted phrase that equals a cluster query is attributed to it. |
| Sep 30 | Audit fixes and follow-ups | Up to five queries per cluster in the writer’s brief; rank reads honour the rank scope and name their lane; a cancelled research run’s competitor artifact is admitted; a stale cluster revision resolves through its key; an unreadable resume ledger yields with nothing purchased. |
| Sep 30 | The operator’s settings | Cluster queries join the seed vocabulary; the strategy’s rows reach the target-page fence through their page URL; the seed field carries provenance; the review reads the operator context and may ask the operator to examine a setting. |
Table 1. Algorithm and functionality advances shipped September 25 to 30, 2026
Two of these advances merit emphasis because they closed loops that had been open since the architecture shipped. Until September 30, the topic lanes that choose what to write had never read the strategy: every writer’s receipt in the window carried the disposition “strategy context unmapped,” and no published article carried a cluster attribution.[9] The strategy governed page ownership, priorities and reporting, but the content the system produced most often was chosen from cached opportunity rows and the operator’s seed terms. The second change is the subject of its own section below: the operator’s settings entered the review as evidence.
Materials and Method
The population is every client property that ran the strategy program in Full Autopilot on September 30, 2026: eight properties, described here by vertical and state under the anonymization policy of this index, under which no property, domain or client is named. Site B is the founding note’s local service client in southern New Hampshire.[1] The others are a heating and cooling contractor in Ohio (Site C), a tree service in Kansas (Site D), a flooring installer in Kansas (Site E), a voice-automation software company (Site F), a healthcare-marketing platform (Site G), an AI voice-agent agency in Florida (Site H) and a landscape design and build firm in Kansas (Site I). The window is the ninety days from June 30 to September 27, 2026, ending on the last day for which Search Console had reported at capture; data was read on September 30, 2026.[9]
Two instruments provide the search series, both from the platform’s own reporting ledger rather than from a manual export. The daily capture records single-day Search Console totals for a property from the day the capture began. The trailing-window capture freezes a 28-day Search Console window whenever a report or overview is generated, and reaches further back than the daily capture on the properties that were connected late. Both inherit Search Console’s aggregation, privacy filtering and reporting lag.[3] The measured rank panel is a weekly paid check of a tracked keyword set at the property’s rank scope; we report the latest panel as a level and do not compare panels across a roster or scope change. Strategy versions, review jobs, articles, dispatches and research spend are counted from their own tables.
- Windows. The first 28 days (June 30 to July 27) and the last 28 days (August 31 to September 27) of the window, compared as daily averages. For Site G, whose daily capture began on July 30, the first window is July 30 to August 26. For Sites E, H and I the capture began too late for a non-overlapping first window, and no ratio is reported.
- Impressions are the primary series, as in the founding note; clicks are reported beside them and average position is reported as a diagnostic.
- Weekly charts use ISO weeks from the daily capture with no smoothing; trailing-window charts plot each captured window at its last day.
- Exclusions. The founding note’s current-events property runs a strategy but not in an authoring mode and is excluded. Days before a property was connected to Search Console carry zero-impression rows with no average position; they are excluded from every average and from the days-with-data count.
| Site | Vertical | Daily capture from | Trailing windows | Latest window (impressions · clicks) |
|---|---|---|---|---|
| Site B | local service client, southern New Hampshire | May 25 | 97 captures, Jun 3 to Sep 30 | Aug 31 to Sep 27: 5,851 · 93 |
| Site C | heating and cooling contractor, Ohio | May 25 | 100 captures, Jun 3 to Sep 30 | Aug 31 to Sep 27: 71,076 · 94 |
| Site D | tree service, Kansas | May 25 | 39 captures, Aug 18 to Sep 29 | Aug 30 to Sep 26: 15,740 · 115 |
| Site E | flooring installer, Kansas | Aug 13 | 43 captures, Aug 18 to Sep 30 | Aug 31 to Sep 27: 20,724 · 30 |
| Site F | voice-automation software company, United States | May 25 | 47 captures, Jul 29 to Sep 30 | Aug 31 to Sep 27: 1,022 · 52 |
| Site G | healthcare-marketing platform, United States | Jul 30 | 48 captures, Aug 2 to Sep 30 | Aug 31 to Sep 27: 1,121 · 24 |
| Site H | AI voice-agent agency, Florida | Aug 24 | 22 captures, Sep 7 to Sep 30 | Aug 31 to Sep 27: 1,695 · 21 |
| Site I | landscape design and build firm, Kansas | Sep 11 | 11 captures, Sep 18 to Sep 29 | Aug 30 to Sep 26: 2,772 · 19 |
Table 2. The eight properties and their instruments
Results: Search Visibility
Four properties have a complete daily series across the window. Three of them gained impressions between the first and last 28 days: Site C from 2,026.1 to 2,538.4 per day (1.25×), Site D from 380.3 to 564.5 (1.48×) and Site F from 21.7 to 36.5 (1.68×). Site B held the plateau the founding note described, at 225.7 and then 209.0 per day (0.93×). Average position improved on all four: Site C from 26.5 to 19.8, Site D from 25.0 to 19.4, Site F from 17.4 to 6.8 and Site B from 14.6 to 14.3.[9]
| Site | First window | Impressions / day | Clicks / day | Last window | Impressions / day | Clicks / day | Ratio | Avg. position |
|---|---|---|---|---|---|---|---|---|
| Site B | Jun 30 to Jul 27 | 225.7 | 4.1 | Aug 31 to Sep 27 | 209.0 | 3.3 | 0.93× | 14.6 to 14.3 |
| Site C | Jun 30 to Jul 27 | 2,026.1 | 3.9 | Aug 31 to Sep 27 | 2,538.4 | 3.4 | 1.25× | 26.5 to 19.8 |
| Site D | Jun 30 to Jul 27 | 380.3 | 2.9 | Aug 31 to Sep 27 | 564.5 | 4.2 | 1.48× | 25.0 to 19.4 |
| Site E | capture began too late | n/a | n/a | Aug 31 to Sep 27 | 740.1 | 1.1 | n/a | n/a to 27.0 |
| Site F | Jun 30 to Jul 27 | 21.7 | 1.4 | Aug 31 to Sep 27 | 36.5 | 1.9 | 1.68× | 17.4 to 6.8 |
| Site G | Jul 30 to Aug 26 | 39.8 | 0.5 | Aug 31 to Sep 27 | 40.0 | 0.9 | 1.01× | 37.5 to 19.5 |
| Site H | capture began too late | n/a | n/a | Aug 31 to Sep 27 | 60.5 | 0.8 | n/a | n/a to 7.5 |
| Site I | capture began too late | n/a | n/a | Sep 11 to Sep 27 (17 days covered) | 177.9 | 1.1 | n/a | n/a to 18.2 |
Table 3. Fixed 28-day windows from the daily capture, daily averages
The trailing-window capture corroborates the daily series over a longer reach. Site C’s first captured window, May 6 to Jun 2, held 49,795 impressions; its last, Aug 31 to Sep 27, held 71,076, across 100 captures. Site E, a flooring installer whose Search Console property was young, went from 700 impressions in the window ending Aug 15 to 20,724 in the window ending Sep 27; a new property’s first months are dominated by indexing and should not be read as an effect of the program.
| Site | Days with data | Impressions | Clicks |
|---|---|---|---|
| Site B | 90 | 19,470 | 355 |
| Site C | 90 | 225,291 | 327 |
| Site D | 89 | 45,102 | 338 |
| Site E | 46 | 38,115 | 52 |
| Site F | 90 | 2,663 | 144 |
| Site G | 60 | 2,421 | 43 |
| Site H | 35 | 1,810 | 24 |
| Site I | 17 | 3,025 | 19 |
Table 4. Ninety-day totals from the daily capture, June 30 to September 27, 2026
Clicks are small on every property. The largest ninety-day total is 355 clicks, on Site B (Table 4). These are local service businesses and young software companies, not publishers, and the founding note’s caution applies with more force here: impressions measure visibility, a plateau can mean a market ceiling, and the instrument that moves revenue for a local business is the map grid.[1] Clicks are reported because the platform’s reporting rules require the smaller number to be visible beside the larger one.[13]
Results: Rankings, Output and Review Activity
The measured rank panel is reported as a level. Between September 15 and September 21 the tracked roster settled at thirty keywords on every property, where the first panels had held between seven and seventy-seven, and on September 29 the panel began measuring a local business at its market city, so the first and latest panels of a property do not measure the same thing. The latest country-scoped panels are shown for every property; Site C’s city-scoped panel of September 30 is shown beside its country-scoped panel of September 28 as an illustration of the instrument.
| Site | Panel weeks | Tracked keywords | In top 10 | In top 100 | Median rank of ranked keywords |
|---|---|---|---|---|---|
| Site B | 11 | 30 | 6 | 6 | 1.0 |
| Site C | 12 | 30 | 0 | 1 | 23.0 |
| Site D | 7 | 29 | 2 | 5 | 13.0 |
| Site E | 7 | 30 | 2 | 3 | 7.0 |
| Site F | 10 | 30 | 1 | 1 | 3.0 |
| Site G | 8 | 30 | 0 | 0 | n/a |
| Site H | 3 | 30 | 0 | 0 | n/a |
| Site I | 2 | 30 | 0 | 1 | 18.0 |
Table 5. Latest weekly measured rank panel per property, country scope, September 21 to 29, 2026
Site C at the country and at the city
Country scope, Sep 28: 0 of 30 tracked keywords in the top ten, 1 in the top hundred, median rank 23. Market city, Sep 30: 14 in the top three, 17 in the top ten, 20 in the top hundred, median rank 2. A heating contractor does not compete nationally for “furnace repair”; the city-scoped panel measures it where its customers search. Every reader of rank in the platform now admits measured rows at that scope, and the writer’s brief names the lane a number came from.[12]
Output and review activity are counted from the platform’s own tables over the ninety days ending at the September 30 read, three days past the search window; fifteen of the activations below fall on September 28 to 30. Across the eight properties that span held 170 nightly review jobs, 98 activated versions of which 28 were authored by the review, 57 published articles and 108 writing dispatches of which 56 were verified live on the client’s site. Paid research across the eight properties cost $102.95 in the window; the nightly reviews and the drafts themselves ran on desktop subscriptions at no marginal cost.[9]
| Site | Versions (all time) | Activated in window | Authored by review | Review jobs | Articles published | Dispatches verified live | Tracked keywords | Research spend |
|---|---|---|---|---|---|---|---|---|
| Site B | 27 | 24 | 9 | 33 | 10 | 10 of 19 | 98 | $13.83 |
| Site C | 28 | 26 | 9 | 39 | 15 | 14 of 27 | 76 | $7.42 |
| Site D | 6 | 5 | 0 | 15 | 8 | 7 of 20 | 76 | $19.68 |
| Site E | 9 | 8 | 0 | 22 | 9 | 9 of 10 | 35 | $11.13 |
| Site F | 5 | 2 | 0 | 14 | 7 | 7 of 12 | 31 | $7.21 |
| Site G | 31 | 29 | 10 | 45 | 6 | 7 of 16 | 26 | $12.81 |
| Site H | 8 | 3 | 0 | 2 | 2 | 2 of 4 | 44 | $21.42 |
| Site I | 1 | 1 | 0 | 0 | 0 | 0 of 0 | 63 | $9.45 |
Table 6. Strategy, review, publication and spend activity in the ninety days ending September 30, 2026
Two patterns in Table 6 are results in their own right. Site C, Site B and Site G each accumulated more than twenty versions, and before the working rules of September 30 most of those versions were identical to their predecessors: the review was re-reading the version it had just authored. The rules that closed that loop are described below. And the dispatch column shows the publication lane’s verified ratio: 56 of 108 dispatches were verified live, with the remainder cancelled by a later retry, failed at the desktop, timed out, or committed and awaiting deployment. A dispatch that fails can be re-issued by a single operator action; nothing in the period withheld a finished post from publication.[12]
Site C provides the most complete record: 28 versions, 26 activations and 39 review jobs in the window, 15 articles published, 14 of 27 dispatches verified live, and $7.42 of paid research. It is also the property on which the topic-selection replay showed the largest effect of the September advances, from two candidates to twelve.
Edge Cases and Verified Corrections
Every change in the period was reviewed adversarially before it merged: one or more finders read the diff for defects and a separate refuter tried to disprove each finding by executing the real modules. Findings that survived were fixed before the version shipped. The cases below are the ones that generalize. Each names a class of error an autonomous system can make quietly. The eight code corrections are pinned by executed tests, and the two cases that concern interpretation are handled as reporting rules of this paper.[10][11]
Review rotation and one review per evidence state
Observed. Before the working rules, the same three projects were reviewed every night and twice on some mornings, and sixty of sixty reviews ended in an activation, most of the resulting versions identical to the last: the review was reading the version it had just authored and proposing the same amendments.
Correction. Projects rotate by last review, a review runs once per evidence state, and a version is written only when an amendment changes the strategy. The correction was not a limit on reviews; every eligible project still takes its turn each night.
Status. Shipped in the period and pinned by executed contract tests.
Fence facts stay off the emitted candidate
Observed. When the strategy’s page-keyword rows first reached the target-page fence through a synthesized page URL, that URL also left the loader on the candidate itself. The run lane excludes any candidate whose URL a recent run already used, so a strategy row on a page that a prior run had written about would have been dropped silently as already seen.
Correction. The URL is a fence-only fact and the emitted candidate shape is byte-identical to the shape before the change. The adversarial review identified the defect by executing the loader before the version shipped.
Status. Shipped in the period and pinned by an executed contract test.
Strategy rows reach the target-page fence
Observed. Activation seeds a page-keyword row per cluster query with the owning page’s path, while every authoring pipeline lists its target pages as full URLs. A text fence comparing the two could never match, so none of the 380 strategy rows across the properties had reached a topic chooser since the rows were introduced.
Correction. The strategy’s rows reach the fence through the URL their page has on the site. The replay that identified the defect also measured the repair: one property’s candidate list grew from two rows to twelve, eight of them strategy rows.
Status. Shipped in the period and pinned by an executed contract test.
A stable input digest
Observed. The operator-context block first ordered pipelines by their last update. The writer bumps that column on every run, so a two-pipeline property could render its block in a different order at two points in the same day, moving the review’s input digest with no setting changed and defeating the same-day reuse rule.
Correction. Pipelines are ordered by name. An executed test bumps the timestamp and asserts identical bytes.
Status. Shipped in the period and pinned by an executed contract test.
Competitors are the operator’s pinned domains
Observed. The grounding module’s third-party roster names large technology and retail brands so that a research brief can recognize a keyword about a vendor. Reused without change in the review, it labelled ordinary seed terms that contained such a word as naming another business.
Correction. The review block recognizes only the domains the operator pinned as competitors.
Status. Shipped in the period and pinned by an executed contract test.
Measured rank at the rank scope
Observed. The writer’s brief could state that a site ranked ninth for a phrase while the next line stated that Search Console had never seen it. The rank came from a daily cache echo at the country level.
Correction. Rank reads honour the project’s rank scope, prefer a measured row within ten days over a newer echo, and name their lane.
Status. Shipped in the period and pinned by an executed contract test.
A scope change is a change of instrument
Observed. On September 30 one property’s tracked keywords were measured at its market city for the first time. Two days earlier, at the country scope, none of its thirty tracked keywords ranked in the top ten; at the city, seventeen did.
Rule. Neither panel is wrong; they answer different questions. A comparison across a scope change would be a claim about the instrument, so this paper reports the two panels side by side and does not compare them.
Status. Reporting rule of this paper.
The quarterly clock anchored on completed research
Observed. The quarterly research refresh was anchored on the strategy’s activation timestamp. Under nightly activations the anchor moved every night, so no living strategy could ever be ninety days old.
Correction. The clock is anchored on the newest completed research run.
Status. Shipped in the period and pinned by an executed contract test.
Operator items persist across nights
Observed. A review that asked the operator to examine a setting would ask again the next night, under a new item key, because it could not see that its own earlier item was still open.
Correction. The block lists the open operator items, the prompt reuses their keys, and the applier drops a duplicate key.
Status. Shipped in the period and pinned by an executed contract test.
Unmeasured days are excluded
Observed. The daily capture writes a row for every calendar day, including days before a property was connected to Search Console; those rows carry zero impressions and no average position. On four properties the connection came weeks into the window, and treating those zeros as measured traffic would produce a spurious increase.
Rule. This paper reports connection dates, counts only measured days as days with data, and compares only windows the instrument covered.
Status. Method of this paper.
The pattern across these cases is the one the compiler paper predicted: the model was rarely the failing component. A text fence, a sort key, a roster reused out of its purpose, a clock anchored on the wrong event, and an instrument read at the wrong scope each produced a confident wrong answer that no prompt could have corrected. Human-AI design guidance and the generative-AI risk profile both ask for traceability and legible correction paths; in this system those are executed tests over the real modules, and a contract test is never deleted, only re-pinned with a dated note.[7][8]
Operator Configuration as Review Evidence
The most consequential advance of the period concerns the settings a person types rather than the evidence plane. Each content pipeline carries a seed field: a list of terms that both guards the vertical, so that an automotive keyword cannot reach a heating contractor’s blog, and expresses the topical focus. On three properties that field held ten loose phrases written months earlier, some of them the names of other businesses. The topic loader admitted only candidates that matched those phrases, so on Site C one of sixty-six strategy rows could pass, and no review could see why the content had drifted from the strategy, because no review read the field.[11]
The setting
An operator types seed terms into the pipeline, or research setup fills them in; since September the field carries a provenance stamp saying which.
The vocabulary
Topic candidates pass the seed fence when they name either an operator seed term or an active cluster query; the fence only widens.
The review
The nightly review reads the seed terms beside the clusters and says which term no cluster holds or which names another business.
The ask
A disagreement becomes one manual-channel backlog item marked for the operator; the review never rewrites the setting and the client never sees the item.
The advance had four parts, shipped as four versions on September 30. The active strategy’s cluster queries joined the seed vocabulary, so a candidate that names one passes the fence even when the operator’s seed does not; the fence only widens, and an empty or brand-only seed stays open exactly as before. The strategy’s own page-keyword rows, which carry a page path while every pipeline lists its target pages as URLs, reach the target-page fence through the URL their page has on the site, as a fence-only fact. The seed field gained a provenance stamp, written only when the seed changes, that says whether a person set it or research filled it in. And the nightly review gained the operator-context block: each seed term with the cluster that holds it, or the fact that none does, or the competitor it names; the banned topics; the writer’s guidelines; and the client summary with its version.
What the review may do with a disagreement is deliberately narrow. It may propose one manual-channel backlog item per issue, marked for the operator, with a title that says what to examine and why. The applier keeps that marker and forces the manual channel, so the item can never appear on a client’s list. No verb rewrites the seed, the guidelines or the summary. This is the human-governance boundary the compiler paper described, applied to the place where it was most needed: the system widens what it considers, names what it disagrees with, and asks.[2][7]
Discussion
The content-only pipeline optimized a single act, writing the next post, against a cache of evidence. The living strategy optimizes a policy: which intents the site should own, which pages own them, which queries belong to which cluster, what to write next and why, and when a claim about the site is too old to act on. The nightly review makes that policy self-updating within a closed vocabulary, and the projections make one version the authority for every lane, including the client’s. To our knowledge no published system recompiles a search strategy nightly from combined first-party, paid, local, technical and site evidence and hands each agent a bounded slice of it; we make the claim of novelty cautiously and would welcome a counter-example.
What the ninety days show is that the system runs as designed at fleet scale. Eight properties took their nightly turns; 28 versions were authored by the review; the review’s input digest coalesced duplicate work; a human draft was never overwritten; and the client portal showed only client-facing items. They also show that four loops of the intended design were closed only in the last week of the period: the topic lanes had not read the strategy, the operator’s settings had been invisible to the review, the quarterly research clock had not come due, and the rank instrument had measured a local business at the wrong scope. Each was identified by reading production rows against the code rather than by a model reporting its own success, and each is now closed.
The visibility results are consistent with the founding observation and no stronger. Three full-span properties gained impressions per day between windows; one held a plateau that its local grid explains; Site G, compared over a shifted first window, was level; three connected too late to compare. Seasonality is confounded on Site C in particular, whose late-summer rise coincides with the heating season. The next study is the one the compiler paper specified: a declared treatment date for a property, a prespecified counterfactual and a lagged comparison, run on properties whose daily capture covers the whole design.[4]
Safeguards for the site’s interest
A system that authors its own policy needs a boundary that the model cannot override. Here the boundary is structural. Amendments live in a closed vocabulary. An owner-written client summary survives every report. The automation caps its own additions below top priority so that a person decides what is urgent. A review that disagrees with a setting asks the operator instead of changing it. And every fence in the content lane may only widen: nothing added in the period can make a post less likely to publish.
Limitations and Future Work
The compiler paper asked for three linked studies: a paired context-quality replay, an orchestration and fault-injection ablation, and a longitudinal study of search outcomes.[2] This paper contributes to the third as a descriptive record and to the first in a narrow form: the topic-loader replays hold model, tools and project snapshot constant and compare the former and current candidate lists on real rows, without a model call. The fault-injection matrix remains to be run as a study, although several of its rows, a stale hash, a duplicate delivery and an unreadable ledger, appeared as production corrections in the period.
- Eight properties in one agency’s portfolio constitute a fleet record, not a sample; no control group exists and no treatment date was declared before the window.
- Search Console figures carry the provider’s aggregation, privacy thresholds and reporting lag; the platform’s ledger records them as reported and never back-fills a day.
- Four properties’ daily capture began inside the window; three of their comparisons are withheld and the fourth uses a shifted first window, and their trailing windows begin at a property’s first indexed weeks.
- The measured rank panel changed roster between September 15 and September 21 and scope on September 29; only the latest panel is reported, as a level.
- Seasonal demand is confounded with the program on at least one property; impressions are visibility, not revenue; clicks are small on every property.
- The September advances shipped in the last week of the window; their effect on what the system writes is deferred to a later report, and the first cluster-attributed articles had not yet published at capture.
- Ledgers and source are private; the series in this paper are published as a curated data module with windows and instruments named, and the figures can be recomputed from it.
Conclusion
A strategy that is recompiled every night from every source a platform holds, amended in a closed vocabulary, and projected as one authority to every agent and to the client is now a system in operation rather than a design. Across eight properties over ninety days it authored 28 of its own versions, published 57 articles, and kept every human decision it was designed to keep. The search results are consistent with the program’s founding observation and are reported as observations.
The principal advances of the period are the loops it closed. Topic selection now reads the strategy, the operator’s configuration is evidence, the quarterly clock is anchored on completed research, and rank is measured where the business competes. The discipline that produced those advances, executed tests over real rows and adversarial review of every change, is the practice this paper recommends carrying unchanged into the causal study specified above.
Data Availability and Document Record
The series plotted and tabulated in this paper are published as a curated data module in which every window, instrument and capture date is named; every chart and every results table is generated from that module and can be recomputed from it; the schematic figures, the change log and the document record are not derived from it. The platform ledgers and source are private and are cited by version and by executed contract suite.[9][12]
| Field | Value |
|---|---|
| Document type | Technical report and longitudinal case study |
| Series and version | Zyan Research, version 1.0 |
| Published | September 30, 2026 |
| Observation window | Fixed-window comparisons and ninety-day totals: June 30 to September 27, 2026; plotted series: the week of June 1 (weekly) and June 3 (trailing windows) through September 30, 2026; activity counts: the ninety days ending September 30, 2026 |
| Population | Eight client properties in Full Autopilot, anonymized by vertical and location |
| Instruments | Search Console daily capture and trailing 28-day windows; weekly measured rank panel; platform ledgers for versions, reviews, articles, dispatches and research spend |
| Data read | September 30, 2026 |
| Companion publications | The Strategy Compiler: A Closed-Loop Context Architecture for Autonomous SEO Agents; Founding Data: What 12x Impressions Growth Actually Looks Like |
| Suggested citation | Zyan Labs (2026). The Living Strategy: Ninety Days of Autonomous SEO Across Eight Properties. Zyan Research, September 30, 2026. |
| Revision history | Version 1.0, September 30, 2026: initial publication |
Table 7. Document record
