Publication

The Living Strategy: Ninety Days of Autonomous SEO Across Eight Properties

This paper documents ninety days of operation of an autonomous SEO system whose strategy is recompiled every night from search, analytics, rank, competitor, local, technical and site evidence and delivered to specialized agents as bounded context. It advances the earlier content-only pipeline and its founding case study in three respects: the strategy is a versioned authority that the agents read rather than a prompt they carry; a nightly review proposes amendments in a closed vocabulary and authors a new version only when an amendment changes the strategy; and, from late September 2026, the operator’s own configuration enters that review as evidence. We document the algorithm and functionality advances shipped in the period, the edge cases identified by adversarial review together with their verified corrections, and the windowed search results of eight client properties. The results are reported descriptively under a stated claim discipline. Three of the four properties with a complete daily series gained impressions between the first and last 28-day windows; one held an established plateau; a fifth, compared over a shifted first window, was level; three properties connected too late for a comparison. No causal effect is claimed, and the design of the causal study that should follow is stated.

Keywords: autonomous search engine optimization; strategy compilation; nightly review; closed amendment vocabulary; context projection; operator configuration; longitudinal case study.

Scope and Claims

Our earlier publication described an architecture in which search, analytics, rank, local, technical and site evidence compile into a versioned strategy that constrains specialized agents.[2] It ended with an evaluation protocol and a boundary: no controlled study had yet measured the architecture, and the founding growth observation predated it.[1] This paper is the first longitudinal record against that protocol. It covers the ninety days from June 30 to September 27, 2026, on the eight client properties that ran the system in Full Autopilot, the mode in which the platform authors and activates strategy versions without a person approving each one.

The claims are bounded by design. We show that a strategy can be recompiled every night from every source the platform holds, that its amendments are bounded and auditable, that the same version reaches topic selection, the writer’s brief, the page and technical queues and the client portal, and that the operator’s own configuration now enters the review as evidence rather than as an invisible constraint. We show the search visibility of the eight properties over the ninety-day window, with each 28-day comparison window named. We do not claim that the architecture caused the visibility: a causal claim requires a declared treatment date and a prespecified counterfactual, and the design that supplies both is stated in the closing sections.[4]

Figure 1 shows the combined Search Console impressions of the eight properties over the ninety-day window, week by week, as the daily capture recorded them: 337,897 impressions and 1,302 clicks between June 30 and September 27, 2026. The roster and instrument facts that qualify the curve are stated in the caption and in Materials and Method.[9]

Hover the curve for weekly totals
010K20K30K40K50KJun 29Jul 13Jul 27Aug 10Aug 24Sep 7Sep 21
Figure 1. Search Console impressions across the eight properties combined, weekly totals from the daily capture, week of June 29 through week of September 21, 2026, spanning the ninety-day window. The series rises from 20,148 impressions in the first week to a peak of 41,016 in the week of August 24 and closes at 23,165. Two facts qualify the curve. Four properties connected inside the span (Site G from July 30, Site E from August 13, Site H from August 24, Site I from September 11) and contribute nothing before their connection, so part of the rise is properties joining the record; the four properties with a complete series (Sites B, C, D and F) sum to 73,453 impressions over the first four weeks and 93,190 over the last four, a ratio of 1.27, against 1.63 for all eight. The decline through September appears on the complete-series properties as well and is recorded as captured, without smoothing; Site D’s capture ends September 26, so its final week carries six days.

Summary of Contributions

The contributions of the period are listed in the order in which the paper documents them. Each is traceable to a versioned change, an executed contract test or a ledger read cited in the sources.[9][10][11][12]

  • A longitudinal record against the evaluation protocol of the compiler paper: eight properties over ninety days, with every window and instrument named.
  • Working rules for the nightly review, namely rotation by last review, one review per evidence state, and the rule that a version is a change, which ended the re-review loop in which sixty of sixty reviews had produced an activation.
  • Strategy-aware topic selection. All three topic lanes read the active clusters and the open article backlog, the strategy’s own page-keyword rows enter the candidate list, and an accepted search phrase is attributed to the cluster whose query it matches.
  • Operator configuration as review evidence, with a bounded path by which the review may ask the operator to examine a setting and no path by which it may rewrite one.
  • Rank measured at the market city. Every reader of rank admits measured rows at the property’s rank scope, and the writer’s brief names the lane a number came from.
  • A quarterly research clock anchored on the newest completed research run rather than on an activation timestamp that every nightly version re-stamped.
  • A verification discipline, adversarial review of every change with execution of the real modules, which identified the eight defects documented below and pinned each correction with an executed test.
  • A recomputable data module from which the search, rank and activity series plotted and tabulated in this paper are generated.

Background: From a Content Pipeline to a Living Strategy

The founding case study documented a feed-forward content pipeline: collect signals, choose a topic, draft, publish, repeat.[1] Its founding property grew more than twelvefold in impressions and its local counterpart reached a plateau within six weeks, and the note argued that the plateau, not the curve, was the more instructive result. Both properties ran the same kind of loop: the model chose a topic from cached evidence, wrote against a brief, and the outcome was measured afterward. Nothing in the loop decided what the site should be about next quarter, which pages should own which search intents, or when a piece of evidence was too old to act on.

The strategy compiler introduced that missing layer as a versioned authority.[2] The system described here is that architecture in operation, with one addition that turned a compiled artifact into a living one: a nightly review that reads the strategy against fresh evidence and proposes amendments in a closed vocabulary, and, under Full Autopilot, authors and activates the next version itself. The adjective is used in its literal sense: across the eight properties, 98 versions were activated in the ninety days ending September 30, 2026, and 28 of them were authored by the review rather than by research or a person.[9]

  1. 01

    Observe

    Search Console queries and pages, analytics sessions, paid rank and keyword data, competitor pages, local grid scans, crawl and index state, page speed, and the live site inventory.

  2. 02

    Qualify

    Each observation keeps its window, scope, method and source; a measured rank outranks a cached echo on the same day; a missing source stays absent rather than zero.

  3. 03

    Compile

    The active strategy binds clusters, owning pages, priorities and a typed backlog under one content hash; activation projects page ownership and tracked terms into operational tables.

  4. 04

    Review nightly

    A desktop turn assembles the strategy, the evidence blocks, the live state and, since September, the operator context; it proposes bounded amendments in a closed vocabulary.

  5. 05

    Author and activate

    Under Full Autopilot, accepted amendments become the next version only when they change something; identical proposals resolve without a version.

  6. 06

    Project

    Topic selection, the writer’s brief, page and technical queues, and the client portal read the same active version; each lane receives its own slice.

  7. 07

    Measure

    Published work, rank checks, search windows and the impact journal return as evidence for the next review under their own clocks.

The loop closes every night. Nothing in it rewrites a human decision: an owner-written summary, a manual draft and a protected page survive every automatic version.
Figure 2. The living strategy as it ran across the eight properties in September 2026. The original content pipeline occupied stages 06 and 07 only; the strategy layer above it is what this paper documents.

Three differences from the content-only pipeline matter for what follows. First, the agents no longer carry the strategy in a prompt; they read a projection of the active version, so two agents working the same property cannot act on different versions of the strategy. Second, adaptation is a version, with a diff, a provenance stamp and an input digest, not an edit to an instruction. Third, the review sees evidence under the same eligibility rules the reports use: a measured rank outranks a cached echo on the same day, a missing source stays absent, and a comparison names its window.[13]

System Description: Evidence Sources and the Decision Path

The platform holds seven families of evidence for a property. Search Console contributes queries, pages, impressions, clicks and position by window. Analytics contributes sessions and organic landing pages. A paid data provider contributes keyword volume and difficulty, the site’s ranked keywords, result-page composition and competitor pages, and the weekly measured rank of a tracked keyword set at the property’s rank scope. Geographic grid scans contribute local map visibility from explicit coordinates. Audits contribute crawl, index, page-speed and schema state. A site inventory contributes the live pages, their canonical paths and internal structure. Research runs contribute a keyword universe, entities and competitor gaps. Each family updates on its own clock.[2]

These sources reach a decision through two compilers. The research pipeline compiles a strategy draft in named stages and validates its synthesis against evidence identifiers. The nightly review compiles a bounded input: nine data blocks serialized deterministically, hashed into an input digest, followed by two prose blocks that carry facts only. The reviewer, a language model running as a desktop turn at no marginal cost, returns findings typed as support, contradiction, gap or amendment, and an amendment may only add or remove a cluster query, set a charter or cluster field, retire a cluster, or add, re-prioritize or drop a backlog item. Anything outside that vocabulary is discarded before it can be applied.[10]

  • strategy

    The active version, its backlog (with any automatic priority cap) and the last diff.

  • gsc_clusters

    Each cluster’s own queries plus the strongest fresh queries no cluster holds.

  • rank

    The measured rank panel at the project’s rank scope, with the cache echo kept apart.

  • competitor_gap

    The research run’s competitor artifact, including a run cancelled at activation.

  • pagespeed · crawl_index

    Technical state from the latest audits and index inspections.

  • impact_journal · monthly_snapshot

    Measured before-and-after windows for shipped changes and the monthly figures.

  • algorithm_updates

    Dated search-engine updates inside the comparison window.

  • operator_context

    Seed terms with the cluster that holds each, banned topics, writer guidelines and the client summary, as facts and days.

  • live_state

    The platform’s own derived reading of the strategy against the evidence it already holds.

Figure 3. What the nightly review reads. The nine data blocks are serialized deterministically and hashed into an input digest; the two prose blocks carry facts and dates only, so a re-assembly over the same facts produces the same bytes and the same digest.

The design follows a principle we have found more durable than any particular prompt: accurate context outperforms restrictions. Long-context studies show that a model does not reliably use the right fact when it sits in the middle of a long input, and our own earlier incident showed a correct file left unread because an index omitted it.[5][14] The review therefore receives a small, complete dossier with stable identifiers rather than the whole client state, and it is asked for decisions inside a vocabulary the applier can validate. Workflows carry the predictable steps; the model is asked only where semantic judgment is necessary.[6]

The projections are equally bounded. Topic selection sees the active clusters, up to six queries each with the owning page, and the open article backlog. The writer’s brief carries the accepted search phrase, the cluster it serves, up to five stored queries per cluster, the site’s standing on that phrase with its instrument named, and the angles recent posts already used. The client portal derives its next and shipped lists from client-facing channels only. The strategy is a single authority with several projections, and no projection is the whole.

Algorithm and Functionality Advances in the Period

The period’s algorithm and functionality advances were shipped as separate, reviewed versions between September 25 and September 30, 2026, each measured against the autoblog doctrine that governs the content lane: context and wording only, no new validator, gate, hold, skip, retry or tool restriction, and nothing that makes a post less likely to publish.[12] The advances are summarized in Table 1; the sections that follow document their effects.

DateChangeWhat it does
Sep 25Automatic authored changesThe nightly review may author and activate the next strategy version under Full Autopilot; a human draft is never overwritten; an owner’s explicit top priority is never claimed by the automation.
Sep 29–30Rank scopeThe weekly rank panel measures a local business at its market city rather than at the country level; every reader of rank admits measured rows at that scope and keeps the daily cache echo apart.
Sep 30Working rules of the reviewProjects rotate by last review; a review never re-reviews the version it just authored; amendments that already hold resolve their findings without a new version; each cluster’s own queries and the strongest unclustered queries appear in the input.
Sep 30The brief is built on the search phraseThe first dispatch of a queued headline carries the phrase the topic was accepted for, and the writer is told why that phrase belongs in the title, heading, slug and description.
Sep 30Nightly cap of eightEvery authoring project takes its turn each night instead of one in three; the turns are desktop time, not paid calls.
Sep 30The quarterly research clockDeep research comes due ninety days after the newest completed run, not after an activation timestamp that every nightly version re-stamped.
Sep 30Owner-written summaries keptA report attach never overwrites a client summary the owner wrote; the portal’s next and shipped lists carry client-facing channels only, in a deterministic order.
Sep 30Topic selection reads the strategyAll three topic lanes see the active clusters and the open article backlog; the strategy’s own page-keyword rows enter the candidate list through the existing fences; an accepted phrase that equals a cluster query is attributed to it.
Sep 30Audit fixes and follow-upsUp to five queries per cluster in the writer’s brief; rank reads honour the rank scope and name their lane; a cancelled research run’s competitor artifact is admitted; a stale cluster revision resolves through its key; an unreadable resume ledger yields with nothing purchased.
Sep 30The operator’s settingsCluster queries join the seed vocabulary; the strategy’s rows reach the target-page fence through their page URL; the seed field carries provenance; the review reads the operator context and may ask the operator to examine a setting.

Table 1. Algorithm and functionality advances shipped September 25 to 30, 2026

Two of these advances merit emphasis because they closed loops that had been open since the architecture shipped. Until September 30, the topic lanes that choose what to write had never read the strategy: every writer’s receipt in the window carried the disposition “strategy context unmapped,” and no published article carried a cluster attribution.[9] The strategy governed page ownership, priorities and reporting, but the content the system produced most often was chosen from cached opportunity rows and the operator’s seed terms. The second change is the subject of its own section below: the operator’s settings entered the review as evidence.

Materials and Method

The population is every client property that ran the strategy program in Full Autopilot on September 30, 2026: eight properties, described here by vertical and state under the anonymization policy of this index, under which no property, domain or client is named. Site B is the founding note’s local service client in southern New Hampshire.[1] The others are a heating and cooling contractor in Ohio (Site C), a tree service in Kansas (Site D), a flooring installer in Kansas (Site E), a voice-automation software company (Site F), a healthcare-marketing platform (Site G), an AI voice-agent agency in Florida (Site H) and a landscape design and build firm in Kansas (Site I). The window is the ninety days from June 30 to September 27, 2026, ending on the last day for which Search Console had reported at capture; data was read on September 30, 2026.[9]

Two instruments provide the search series, both from the platform’s own reporting ledger rather than from a manual export. The daily capture records single-day Search Console totals for a property from the day the capture began. The trailing-window capture freezes a 28-day Search Console window whenever a report or overview is generated, and reaches further back than the daily capture on the properties that were connected late. Both inherit Search Console’s aggregation, privacy filtering and reporting lag.[3] The measured rank panel is a weekly paid check of a tracked keyword set at the property’s rank scope; we report the latest panel as a level and do not compare panels across a roster or scope change. Strategy versions, review jobs, articles, dispatches and research spend are counted from their own tables.

  • Windows. The first 28 days (June 30 to July 27) and the last 28 days (August 31 to September 27) of the window, compared as daily averages. For Site G, whose daily capture began on July 30, the first window is July 30 to August 26. For Sites E, H and I the capture began too late for a non-overlapping first window, and no ratio is reported.
  • Impressions are the primary series, as in the founding note; clicks are reported beside them and average position is reported as a diagnostic.
  • Weekly charts use ISO weeks from the daily capture with no smoothing; trailing-window charts plot each captured window at its last day.
  • Exclusions. The founding note’s current-events property runs a strategy but not in an authoring mode and is excluded. Days before a property was connected to Search Console carry zero-impression rows with no average position; they are excluded from every average and from the days-with-data count.
SiteVerticalDaily capture fromTrailing windowsLatest window (impressions · clicks)
Site Blocal service client, southern New HampshireMay 2597 captures, Jun 3 to Sep 30Aug 31 to Sep 27: 5,851 · 93
Site Cheating and cooling contractor, OhioMay 25100 captures, Jun 3 to Sep 30Aug 31 to Sep 27: 71,076 · 94
Site Dtree service, KansasMay 2539 captures, Aug 18 to Sep 29Aug 30 to Sep 26: 15,740 · 115
Site Eflooring installer, KansasAug 1343 captures, Aug 18 to Sep 30Aug 31 to Sep 27: 20,724 · 30
Site Fvoice-automation software company, United StatesMay 2547 captures, Jul 29 to Sep 30Aug 31 to Sep 27: 1,022 · 52
Site Ghealthcare-marketing platform, United StatesJul 3048 captures, Aug 2 to Sep 30Aug 31 to Sep 27: 1,121 · 24
Site HAI voice-agent agency, FloridaAug 2422 captures, Sep 7 to Sep 30Aug 31 to Sep 27: 1,695 · 21
Site Ilandscape design and build firm, KansasSep 1111 captures, Sep 18 to Sep 29Aug 30 to Sep 26: 2,772 · 19

Table 2. The eight properties and their instruments

Results: Search Visibility

Four properties have a complete daily series across the window. Three of them gained impressions between the first and last 28 days: Site C from 2,026.1 to 2,538.4 per day (1.25×), Site D from 380.3 to 564.5 (1.48×) and Site F from 21.7 to 36.5 (1.68×). Site B held the plateau the founding note described, at 225.7 and then 209.0 per day (0.93×). Average position improved on all four: Site C from 26.5 to 19.8, Site D from 25.0 to 19.4, Site F from 17.4 to 6.8 and Site B from 14.6 to 14.3.[9]

SiteFirst windowImpressions / dayClicks / dayLast windowImpressions / dayClicks / dayRatioAvg. position
Site BJun 30 to Jul 27225.74.1Aug 31 to Sep 27209.03.30.93×14.6 to 14.3
Site CJun 30 to Jul 272,026.13.9Aug 31 to Sep 272,538.43.41.25×26.5 to 19.8
Site DJun 30 to Jul 27380.32.9Aug 31 to Sep 27564.54.21.48×25.0 to 19.4
Site Ecapture began too laten/an/aAug 31 to Sep 27740.11.1n/an/a to 27.0
Site FJun 30 to Jul 2721.71.4Aug 31 to Sep 2736.51.91.68×17.4 to 6.8
Site GJul 30 to Aug 2639.80.5Aug 31 to Sep 2740.00.91.01×37.5 to 19.5
Site Hcapture began too laten/an/aAug 31 to Sep 2760.50.8n/an/a to 7.5
Site Icapture began too laten/an/aSep 11 to Sep 27 (17 days covered)177.91.1n/an/a to 18.2

Table 3. Fixed 28-day windows from the daily capture, daily averages

Hover the curve for weekly totals
010K20K30KJun 1Jun 22Jul 13Aug 3Aug 24Sep 21
Figure 4. Site C, a heating and cooling contractor in Ohio: weekly Search Console impressions from the daily capture, week of June 1 through week of September 21, 2026. Twenty-one of the property’s twenty-eight strategy versions were activated inside this span, as were all fifteen articles published in the ninety-day window; the remaining versions were activated September 28 to 30 or never activated. The late-summer rise coincides with the seasonal demand for heating work and cannot be separated from it here.
Hover the curve for weekly totals
01K2K3K4K5KJun 1Jun 22Jul 13Aug 3Aug 24Sep 14
Figure 5. Site D, a tree service in Kansas: weekly Search Console impressions, week of June 1 through week of September 14, 2026; the property’s daily capture ends September 26, so the partial final week is omitted. Impressions per day rose from 380.3 in the first window to 564.5 in the last while average position improved from 25.0 to 19.4.
Hover the curve for weekly totals
05001K1.5K2K2.5KJun 1Jun 22Jul 13Aug 3Aug 24Sep 21
Figure 6. Site B, the founding note’s local service client: weekly Search Console impressions, week of June 1 through week of September 21, 2026. The plateau documented in August held through the window; the local grid, not the impressions curve, remains the decision-relevant instrument for this property.
Hover the curve for weekly totals
0100200300400Jun 1Jun 22Jul 13Aug 3Aug 24Sep 21
Figure 7. Site F, a voice-automation software company: weekly Search Console impressions, week of June 1 through week of September 21, 2026. The absolute level is small; the ratio of 1.68× between windows and the move in average position from 17.4 to 6.8 are reported as observations at that scale.

The trailing-window capture corroborates the daily series over a longer reach. Site C’s first captured window, May 6 to Jun 2, held 49,795 impressions; its last, Aug 31 to Sep 27, held 71,076, across 100 captures. Site E, a flooring installer whose Search Console property was young, went from 700 impressions in the window ending Aug 15 to 20,724 in the window ending Sep 27; a new property’s first months are dominated by indexing and should not be read as an effect of the program.

Hover the curve for window totals
025K50K75K100KJun 2Jun 18Jul 2Jul 19Aug 13Aug 28Sep 11Sep 27
Figure 8. Site C: trailing 28-day Search Console windows as captured, June 3 through September 30, 2026, plotted at the window’s last day; 98 of 100 captures are shown. Consecutive captures overlap by up to twenty-seven days, so the curve is smoother than the weekly series by construction; the level rises from roughly 50,000 to roughly 71,000 impressions per window. The captures of June 6 and June 11 are omitted because the ledger recorded them with a range other than twenty-eight days (185,004 and 146,843 impressions, the latter over March 13 to June 10); they are reported here rather than plotted.
Hover the curve for window totals
010K20K30KAug 15Aug 22Aug 30Sep 6Sep 13Sep 20Sep 27
Figure 9. Site E: trailing 28-day Search Console windows as captured, August 18 through September 30, 2026. The series begins at a young property’s first indexed weeks; the rise reflects a property entering the index as much as any work performed on it.
337,897
Impressions across the eight properties, daily capture, June 30 to September 27
Four properties partial; see Table 2.
1,302
Clicks across the eight properties, same window and instrument
3 of 4
Full-span properties with more impressions per day in the last window than the first
SiteDays with dataImpressionsClicks
Site B9019,470355
Site C90225,291327
Site D8945,102338
Site E4638,11552
Site F902,663144
Site G602,42143
Site H351,81024
Site I173,02519

Table 4. Ninety-day totals from the daily capture, June 30 to September 27, 2026

Clicks are small on every property. The largest ninety-day total is 355 clicks, on Site B (Table 4). These are local service businesses and young software companies, not publishers, and the founding note’s caution applies with more force here: impressions measure visibility, a plateau can mean a market ceiling, and the instrument that moves revenue for a local business is the map grid.[1] Clicks are reported because the platform’s reporting rules require the smaller number to be visible beside the larger one.[13]

Results: Rankings, Output and Review Activity

The measured rank panel is reported as a level. Between September 15 and September 21 the tracked roster settled at thirty keywords on every property, where the first panels had held between seven and seventy-seven, and on September 29 the panel began measuring a local business at its market city, so the first and latest panels of a property do not measure the same thing. The latest country-scoped panels are shown for every property; Site C’s city-scoped panel of September 30 is shown beside its country-scoped panel of September 28 as an illustration of the instrument.

SitePanel weeksTracked keywordsIn top 10In top 100Median rank of ranked keywords
Site B1130661.0
Site C12300123.0
Site D7292513.0
Site E730237.0
Site F1030113.0
Site G83000n/a
Site H33000n/a
Site I2300118.0

Table 5. Latest weekly measured rank panel per property, country scope, September 21 to 29, 2026

Site C at the country and at the city

Country scope, Sep 28: 0 of 30 tracked keywords in the top ten, 1 in the top hundred, median rank 23. Market city, Sep 30: 14 in the top three, 17 in the top ten, 20 in the top hundred, median rank 2. A heating contractor does not compete nationally for “furnace repair”; the city-scoped panel measures it where its customers search. Every reader of rank in the platform now admits measured rows at that scope, and the writer’s brief names the lane a number came from.[12]

Output and review activity are counted from the platform’s own tables over the ninety days ending at the September 30 read, three days past the search window; fifteen of the activations below fall on September 28 to 30. Across the eight properties that span held 170 nightly review jobs, 98 activated versions of which 28 were authored by the review, 57 published articles and 108 writing dispatches of which 56 were verified live on the client’s site. Paid research across the eight properties cost $102.95 in the window; the nightly reviews and the drafts themselves ran on desktop subscriptions at no marginal cost.[9]

SiteVersions (all time)Activated in windowAuthored by reviewReview jobsArticles publishedDispatches verified liveTracked keywordsResearch spend
Site B27249331010 of 1998$13.83
Site C28269391514 of 2776$7.42
Site D6501587 of 2076$19.68
Site E9802299 of 1035$11.13
Site F5201477 of 1231$7.21
Site G3129104567 of 1626$12.81
Site H830222 of 444$21.42
Site I110000 of 063$9.45

Table 6. Strategy, review, publication and spend activity in the ninety days ending September 30, 2026

Two patterns in Table 6 are results in their own right. Site C, Site B and Site G each accumulated more than twenty versions, and before the working rules of September 30 most of those versions were identical to their predecessors: the review was re-reading the version it had just authored. The rules that closed that loop are described below. And the dispatch column shows the publication lane’s verified ratio: 56 of 108 dispatches were verified live, with the remainder cancelled by a later retry, failed at the desktop, timed out, or committed and awaiting deployment. A dispatch that fails can be re-issued by a single operator action; nothing in the period withheld a finished post from publication.[12]

Site C provides the most complete record: 28 versions, 26 activations and 39 review jobs in the window, 15 articles published, 14 of 27 dispatches verified live, and $7.42 of paid research. It is also the property on which the topic-selection replay showed the largest effect of the September advances, from two candidates to twelve.

Edge Cases and Verified Corrections

Every change in the period was reviewed adversarially before it merged: one or more finders read the diff for defects and a separate refuter tried to disprove each finding by executing the real modules. Findings that survived were fixed before the version shipped. The cases below are the ones that generalize. Each names a class of error an autonomous system can make quietly. The eight code corrections are pinned by executed tests, and the two cases that concern interpretation are handled as reporting rules of this paper.[10][11]

Review rotation and one review per evidence state

Observed. Before the working rules, the same three projects were reviewed every night and twice on some mornings, and sixty of sixty reviews ended in an activation, most of the resulting versions identical to the last: the review was reading the version it had just authored and proposing the same amendments.

Correction. Projects rotate by last review, a review runs once per evidence state, and a version is written only when an amendment changes the strategy. The correction was not a limit on reviews; every eligible project still takes its turn each night.

Status. Shipped in the period and pinned by executed contract tests.

Fence facts stay off the emitted candidate

Observed. When the strategy’s page-keyword rows first reached the target-page fence through a synthesized page URL, that URL also left the loader on the candidate itself. The run lane excludes any candidate whose URL a recent run already used, so a strategy row on a page that a prior run had written about would have been dropped silently as already seen.

Correction. The URL is a fence-only fact and the emitted candidate shape is byte-identical to the shape before the change. The adversarial review identified the defect by executing the loader before the version shipped.

Status. Shipped in the period and pinned by an executed contract test.

Strategy rows reach the target-page fence

Observed. Activation seeds a page-keyword row per cluster query with the owning page’s path, while every authoring pipeline lists its target pages as full URLs. A text fence comparing the two could never match, so none of the 380 strategy rows across the properties had reached a topic chooser since the rows were introduced.

Correction. The strategy’s rows reach the fence through the URL their page has on the site. The replay that identified the defect also measured the repair: one property’s candidate list grew from two rows to twelve, eight of them strategy rows.

Status. Shipped in the period and pinned by an executed contract test.

A stable input digest

Observed. The operator-context block first ordered pipelines by their last update. The writer bumps that column on every run, so a two-pipeline property could render its block in a different order at two points in the same day, moving the review’s input digest with no setting changed and defeating the same-day reuse rule.

Correction. Pipelines are ordered by name. An executed test bumps the timestamp and asserts identical bytes.

Status. Shipped in the period and pinned by an executed contract test.

Competitors are the operator’s pinned domains

Observed. The grounding module’s third-party roster names large technology and retail brands so that a research brief can recognize a keyword about a vendor. Reused without change in the review, it labelled ordinary seed terms that contained such a word as naming another business.

Correction. The review block recognizes only the domains the operator pinned as competitors.

Status. Shipped in the period and pinned by an executed contract test.

Measured rank at the rank scope

Observed. The writer’s brief could state that a site ranked ninth for a phrase while the next line stated that Search Console had never seen it. The rank came from a daily cache echo at the country level.

Correction. Rank reads honour the project’s rank scope, prefer a measured row within ten days over a newer echo, and name their lane.

Status. Shipped in the period and pinned by an executed contract test.

A scope change is a change of instrument

Observed. On September 30 one property’s tracked keywords were measured at its market city for the first time. Two days earlier, at the country scope, none of its thirty tracked keywords ranked in the top ten; at the city, seventeen did.

Rule. Neither panel is wrong; they answer different questions. A comparison across a scope change would be a claim about the instrument, so this paper reports the two panels side by side and does not compare them.

Status. Reporting rule of this paper.

The quarterly clock anchored on completed research

Observed. The quarterly research refresh was anchored on the strategy’s activation timestamp. Under nightly activations the anchor moved every night, so no living strategy could ever be ninety days old.

Correction. The clock is anchored on the newest completed research run.

Status. Shipped in the period and pinned by an executed contract test.

Operator items persist across nights

Observed. A review that asked the operator to examine a setting would ask again the next night, under a new item key, because it could not see that its own earlier item was still open.

Correction. The block lists the open operator items, the prompt reuses their keys, and the applier drops a duplicate key.

Status. Shipped in the period and pinned by an executed contract test.

Unmeasured days are excluded

Observed. The daily capture writes a row for every calendar day, including days before a property was connected to Search Console; those rows carry zero impressions and no average position. On four properties the connection came weeks into the window, and treating those zeros as measured traffic would produce a spurious increase.

Rule. This paper reports connection dates, counts only measured days as days with data, and compares only windows the instrument covered.

Status. Method of this paper.

The pattern across these cases is the one the compiler paper predicted: the model was rarely the failing component. A text fence, a sort key, a roster reused out of its purpose, a clock anchored on the wrong event, and an instrument read at the wrong scope each produced a confident wrong answer that no prompt could have corrected. Human-AI design guidance and the generative-AI risk profile both ask for traceability and legible correction paths; in this system those are executed tests over the real modules, and a contract test is never deleted, only re-pinned with a dated note.[7][8]

Operator Configuration as Review Evidence

The most consequential advance of the period concerns the settings a person types rather than the evidence plane. Each content pipeline carries a seed field: a list of terms that both guards the vertical, so that an automotive keyword cannot reach a heating contractor’s blog, and expresses the topical focus. On three properties that field held ten loose phrases written months earlier, some of them the names of other businesses. The topic loader admitted only candidates that matched those phrases, so on Site C one of sixty-six strategy rows could pass, and no review could see why the content had drifted from the strategy, because no review read the field.[11]

  • The setting

    An operator types seed terms into the pipeline, or research setup fills them in; since September the field carries a provenance stamp saying which.

  • The vocabulary

    Topic candidates pass the seed fence when they name either an operator seed term or an active cluster query; the fence only widens.

  • The review

    The nightly review reads the seed terms beside the clusters and says which term no cluster holds or which names another business.

  • The ask

    A disagreement becomes one manual-channel backlog item marked for the operator; the review never rewrites the setting and the client never sees the item.

The operator retains final authority over the setting. The system widens what it considers, names what it disagrees with, and asks; it does not decide for the person who configured the site.
Figure 10. How an operator setting reaches a decision after the September changes. Before them, the seed field alone fenced topic selection on three properties and no review could see it.

The advance had four parts, shipped as four versions on September 30. The active strategy’s cluster queries joined the seed vocabulary, so a candidate that names one passes the fence even when the operator’s seed does not; the fence only widens, and an empty or brand-only seed stays open exactly as before. The strategy’s own page-keyword rows, which carry a page path while every pipeline lists its target pages as URLs, reach the target-page fence through the URL their page has on the site, as a fence-only fact. The seed field gained a provenance stamp, written only when the seed changes, that says whether a person set it or research filled it in. And the nightly review gained the operator-context block: each seed term with the cluster that holds it, or the fact that none does, or the competitor it names; the banned topics; the writer’s guidelines; and the client summary with its version.

What the review may do with a disagreement is deliberately narrow. It may propose one manual-channel backlog item per issue, marked for the operator, with a title that says what to examine and why. The applier keeps that marker and forces the manual channel, so the item can never appear on a client’s list. No verb rewrites the seed, the guidelines or the summary. This is the human-governance boundary the compiler paper described, applied to the place where it was most needed: the system widens what it considers, names what it disagrees with, and asks.[2][7]

Discussion

The content-only pipeline optimized a single act, writing the next post, against a cache of evidence. The living strategy optimizes a policy: which intents the site should own, which pages own them, which queries belong to which cluster, what to write next and why, and when a claim about the site is too old to act on. The nightly review makes that policy self-updating within a closed vocabulary, and the projections make one version the authority for every lane, including the client’s. To our knowledge no published system recompiles a search strategy nightly from combined first-party, paid, local, technical and site evidence and hands each agent a bounded slice of it; we make the claim of novelty cautiously and would welcome a counter-example.

What the ninety days show is that the system runs as designed at fleet scale. Eight properties took their nightly turns; 28 versions were authored by the review; the review’s input digest coalesced duplicate work; a human draft was never overwritten; and the client portal showed only client-facing items. They also show that four loops of the intended design were closed only in the last week of the period: the topic lanes had not read the strategy, the operator’s settings had been invisible to the review, the quarterly research clock had not come due, and the rank instrument had measured a local business at the wrong scope. Each was identified by reading production rows against the code rather than by a model reporting its own success, and each is now closed.

The visibility results are consistent with the founding observation and no stronger. Three full-span properties gained impressions per day between windows; one held a plateau that its local grid explains; Site G, compared over a shifted first window, was level; three connected too late to compare. Seasonality is confounded on Site C in particular, whose late-summer rise coincides with the heating season. The next study is the one the compiler paper specified: a declared treatment date for a property, a prespecified counterfactual and a lagged comparison, run on properties whose daily capture covers the whole design.[4]

Safeguards for the site’s interest

A system that authors its own policy needs a boundary that the model cannot override. Here the boundary is structural. Amendments live in a closed vocabulary. An owner-written client summary survives every report. The automation caps its own additions below top priority so that a person decides what is urgent. A review that disagrees with a setting asks the operator instead of changing it. And every fence in the content lane may only widen: nothing added in the period can make a post less likely to publish.

Limitations and Future Work

The compiler paper asked for three linked studies: a paired context-quality replay, an orchestration and fault-injection ablation, and a longitudinal study of search outcomes.[2] This paper contributes to the third as a descriptive record and to the first in a narrow form: the topic-loader replays hold model, tools and project snapshot constant and compare the former and current candidate lists on real rows, without a model call. The fault-injection matrix remains to be run as a study, although several of its rows, a stale hash, a duplicate delivery and an unreadable ledger, appeared as production corrections in the period.

  • Eight properties in one agency’s portfolio constitute a fleet record, not a sample; no control group exists and no treatment date was declared before the window.
  • Search Console figures carry the provider’s aggregation, privacy thresholds and reporting lag; the platform’s ledger records them as reported and never back-fills a day.
  • Four properties’ daily capture began inside the window; three of their comparisons are withheld and the fourth uses a shifted first window, and their trailing windows begin at a property’s first indexed weeks.
  • The measured rank panel changed roster between September 15 and September 21 and scope on September 29; only the latest panel is reported, as a level.
  • Seasonal demand is confounded with the program on at least one property; impressions are visibility, not revenue; clicks are small on every property.
  • The September advances shipped in the last week of the window; their effect on what the system writes is deferred to a later report, and the first cluster-attributed articles had not yet published at capture.
  • Ledgers and source are private; the series in this paper are published as a curated data module with windows and instruments named, and the figures can be recomputed from it.

Conclusion

A strategy that is recompiled every night from every source a platform holds, amended in a closed vocabulary, and projected as one authority to every agent and to the client is now a system in operation rather than a design. Across eight properties over ninety days it authored 28 of its own versions, published 57 articles, and kept every human decision it was designed to keep. The search results are consistent with the program’s founding observation and are reported as observations.

The principal advances of the period are the loops it closed. Topic selection now reads the strategy, the operator’s configuration is evidence, the quarterly clock is anchored on completed research, and rank is measured where the business competes. The discipline that produced those advances, executed tests over real rows and adversarial review of every change, is the practice this paper recommends carrying unchanged into the causal study specified above.

Data Availability and Document Record

The series plotted and tabulated in this paper are published as a curated data module in which every window, instrument and capture date is named; every chart and every results table is generated from that module and can be recomputed from it; the schematic figures, the change log and the document record are not derived from it. The platform ledgers and source are private and are cited by version and by executed contract suite.[9][12]

FieldValue
Document typeTechnical report and longitudinal case study
Series and versionZyan Research, version 1.0
PublishedSeptember 30, 2026
Observation windowFixed-window comparisons and ninety-day totals: June 30 to September 27, 2026; plotted series: the week of June 1 (weekly) and June 3 (trailing windows) through September 30, 2026; activity counts: the ninety days ending September 30, 2026
PopulationEight client properties in Full Autopilot, anonymized by vertical and location
InstrumentsSearch Console daily capture and trailing 28-day windows; weekly measured rank panel; platform ledgers for versions, reviews, articles, dispatches and research spend
Data readSeptember 30, 2026
Companion publicationsThe Strategy Compiler: A Closed-Loop Context Architecture for Autonomous SEO Agents; Founding Data: What 12x Impressions Growth Actually Looks Like
Suggested citationZyan Labs (2026). The Living Strategy: Ninety Days of Autonomous SEO Across Eight Properties. Zyan Research, September 30, 2026.
Revision historyVersion 1.0, September 30, 2026: initial publication

Table 7. Document record

Sources & Notes

  1. 1.

    Founding Data: What 12x Impressions Growth Actually Looks Like. Zyan Research · September 30, 2026

  2. 2.

    The Strategy Compiler: A Closed-Loop Context Architecture for Autonomous SEO Agents. Zyan Research · September 30, 2026

  3. 3.

    About Search Console data: aggregation, privacy filtering, data limits, and interpretation. Primary platform documentation · Google Search Central · September 30, 2026

  4. 4.

    Brodersen et al. (2015), Inferring causal impact using Bayesian structural time-series models. Peer-reviewed · Annals of Applied Statistics, 9(1), 247–274 · September 30, 2026

  5. 5.

    Liu et al. (2024), Lost in the Middle: How language models use long contexts. Peer-reviewed · Transactions of the ACL, 12, 157–173 · September 30, 2026

  6. 6.

    Anthropic (2024), Building effective agents: workflows, agents, tools, and evaluation. Industry engineering guidance · Anthropic · September 30, 2026

  7. 7.

    Amershi et al. (2019), Guidelines for human-AI interaction. Peer-reviewed · ACM CHI Conference on Human Factors in Computing Systems · September 30, 2026

  8. 8.

    National Institute of Standards and Technology (2024), Artificial Intelligence Risk Management Framework: Generative AI Profile. Government guidance · National Institute of Standards and Technology · September 30, 2026

  9. 9.

    Platform reporting ledger for the eight Full Autopilot properties: single-day Search Console captures, trailing 28-day report windows, the weekly measured rank panel, strategy versions and events, worker jobs, blog dispatches and research spend, read on September 30, 2026. Internal observational record · Zyan Labs · September 30, 2026

  10. 10.

    Nightly review implementation: the shared strategy-review module (block assembly, input digest, closed amendment vocabulary, the operator-context and live-state prose blocks), the automation-policy applier and their executed contract suites. Private implementation evidence · Zyan source and contract-test record · September 30, 2026

  11. 11.

    Topic selection and writer’s brief implementation: the strategy topic context, the candidate loader and its fences, the accepted-search-phrase mint, the rank-scope reads and the seed-provenance module, with their executed contract suites and the replays against exported production rows. Private implementation evidence · Zyan source and contract-test record · September 30, 2026

  12. 12.

    Change record, September 25–30, 2026: SEO Command Center v3.0.97 and v3.0.115 through v3.0.127, research satellite v1.5.39 through v1.5.43, client portal performance function v1.0.9, with the dated owner rulings recorded in the autoblog doctrine. Private implementation evidence · Zyan changelog and doctrine record · September 30, 2026

  13. 13.

    Honesty Engineering: Teaching an Autonomous Platform to Tell the Truth. Zyan Research · September 30, 2026

  14. 14.

    Context Is an Engineering Surface: How AI Agents Read a Website. Zyan Research · September 30, 2026