Coscientific Swarm

Coscientific Swarm

Coscientific Swarm is an architecture for parallel research.

  • Independent workers explore a hypothesis graph.
  • They share findings through a Queen-maintained knowledge corpus.
  • They take work from a prioritized research frontier.

The swarm pursues one global goal until the completion criteria hold, or another explicit stop applies.

Inspiration

I started from Andrej Karpathy's autoresearch. In Karpathy's Autoresearch, an agent edits training code, runs a five-minute experiment, keeps or discards the change, and repeats. Coscientific Swarm adds coordination around that loop: competing hypotheses, parallel labs, shared knowledge, and global scheduling.

I also took inspirations from OpenAI's The Hugging Face incident and the road ahead, 26 August 2026. OpenAI described evaluation agents that used unauthorized message boards to keep discoveries, exchange information, coordinate work, and continue investigations other agents started.

I wanted that channel explicit, scoped, and verifiable. Labs submit results to the Queen. Other labs can then verify or refute a finding, the way papers get checked in academia.

Shared state

The Queen is the only writer of shared state. Workers submit result packages, like papers in academia. The Queen writes the corpus and sets final scheduling priority.

  • The research graph records hypotheses, origins, methods, results, evidence, parent and starting-weight links, and execution history. It answers what was investigated, why, and with what outcome.
  • The knowledge corpus holds reusable findings drawn from that graph. Every substantive entry points at its evidence. It answers what a worker should know before the next investigation.
  • The research frontier holds eligible, unclaimed work in scheduling order. It answers which unresolved question should get capacity next.
flowchart TD
    Goal["Swarm objective"] --> O["Queen"]
    O -->|"Commit records"| Graph[("Research graph")]
    Graph -->|"History and evidence"| O
    O -->|"Publish versioned knowledge"| Corpus[("Knowledge corpus")]
    Corpus -->|"Relevant findings"| O
    O -->|"Admit and reprioritize"| Frontier[("Research frontier")]
    Frontier -->|"Eligible candidates"| Scheduler["Scheduler"]
    O -->|"Run state and resource policy"| Scheduler
    Scheduler -->|"Atomically lease one compatible task"| Context["Context compiler"]
    Goal --> Context
    Graph -->|"Relevant records"| Context
    Corpus -->|"Knowledge snapshot"| Context
    Context --> Lab1["Lab 1"]
    Context --> Lab2["Lab 2"]
    Context --> Lab3["Lab 3"]
    Context --> Lab4["Lab 4"]
    Lab1 -->|"Result and 0-3 proposals"| O
    Lab2 -->|"Result and 0-3 proposals"| O
    Lab3 -->|"Result and 0-3 proposals"| O
    Lab4 -->|"Result and 0-3 proposals"| O
    Goal --> Judge["Goal judge"]
    Graph -->|"Evidence and artifacts"| Judge
    Corpus -->|"Consolidated findings"| Judge
    Judge -->|"Continue or verified completion"| O

Labs return results to the Queen. The Queen updates the graph and corpus. The scheduler leases work from the frontier. The context compiler attaches relevant shared knowledge to each lease. The goal judge changes run state through the Queen.

The Queen is a logical authority and can be implemented as more than one program. Validation, corpus updates, candidate evaluation, and scheduling can be separate, as long as write rights and identifiers stay consistent.

Goal

A goal names the outcome and the conditions for declaring success. Changing the goal mid-research is heavily discouraged, since it voids a lot of scientific findings by the previous labs.

python
class GoalContract:
    """
    Sketch of a research or engineering goal.
    """
    objective: str
    success_criteria: str
    constraints: str
    evidence_requirements: str
    resource_budget: str
    permitted_tools_and_environments: str
    stopping_policy: str
    version: str

Nodes and attempts

A research node is one bounded question. It has to be runnable without the proposer's conversation.

python
class ResearchNode:
    """
    Rough sketch of a research node.
    """
    id: str
    goal_version: str
    question: str
    hypothesis: str
    rationale: str
    scope_and_assumptions: str
    predictions: str
    falsification_or_resolution_criteria: str
    origin_references: str
    related_node_ids: str
    evidence_references: str
    status: str
    notes: str
    conclusion: str
    conclusion_confidence: str
    proposed_followups: list  # 0-3
    scoring_history: str
    provenance: str
    starting_model_id: str  # or weight

Execution attempts belong to the node and are recorded separately. A crashed environment is recorded as an execution failure, and the hypothesis stays open. A retry creates another attempt on the same node.

python
class ResearchAttempt:
    """
    Rough sketch of one execution attempt.
    """
    id: str
    node_id: str
    worker_id: str
    methodology: str
    input_artifacts: str
    environment_version: str
    context_snapshot: str
    lease: str
    status: str
    observations: str
    output_artifacts: str
    resource_usage: str

A conclusion says whether the evidence supports, contradicts, is inconclusive, or is limited in scope. Confidence is an assessment of a claim and its evidence.

Follow-ups and admission

On completion, a worker may propose up to three follow-up hypotheses. Each proposal states a rationale, why it could be useful, confidence, proposed importance, and estimated cost. Zero is valid. Follow-ups exist only when the result suggests them.

A falsified hypothesis can close a branch. The negative finding stays searchable. A negative result can also open a real alternative. The cap of three limits branching at one node. Admission thresholds, deduplication, and budgets still bound the program.

Every proposal is compared with active, completed, deferred, and closed research before it gets a new identity or enters the frontier. Similar wording is a retrieval hint. Equivalence depends on matching datasets, interventions, conditions, and outcomes. The Queen compares those before collapsing two nodes.

  • Equivalent: the existing node stays canonical, and the proposal records a reference to it. A running node keeps its execution. A completed node supplies its recorded result.
  • Partial overlap: the unresolved distinction stays. Compatible context can be merged into the canonical record, or a narrower follow-up can be admitted. A merge still goes through evaluation.
  • New: score importance, usefulness, and cost. Admit, defer, or reject with a rationale.

Intentional replication is a new attempt, or an explicit replication task, on the same hypothesis, with purpose and cost recorded. Ancestry stays acyclic. Equivalence, support, and contradiction links can form a richer graph. Reusing an ancestor leaves the parent-child history acyclic.

Corpus and context

The corpus is a maintained set of reusable findings, limits, methods, artifacts, failed approaches, and open disagreements. A result stays attached to the dataset and conditions that produced it.

python
class KnowledgeEntry:
    """
    Sketch of a corpus entry.
    """
    id: str
    version: str
    statement: str
    scope_and_conditions: str
    status: str
    confidence_assessment: str
    supporting_evidence: str
    contradicting_evidence: str
    source_node_ids: str
    artifact_references: str
    supersedes: str
    created_at: str
    updated_at: str

Status can be provisional, supported in a stated scope, disputed, superseded, or invalidated. Those labels describe status within that scope.

The Queen checks a submitted result against its evidence and the existing corpus, then publishes a versioned update. Raw observations and prior versions stay. A new summary leaves conflicting evidence in place. Workers submit interpretations. The Queen publishes a derived account of the graph. Later labs supply independent verification.

That is the cross-lab path. Lab A submits to the Queen. Lab B receives a snapshot through the context compiler. Labs exchange findings through those two, on read-only snapshots.

Before a task starts, the compiler assembles the goal, hypothesis, relevant ancestry, related experiments, applicable corpus entries, contradictions, tools, and limits. It records which knowledge versions it supplied, so later you can see whether a worker lacked a finding, ignored one, or used a conclusion that was later revised.

New knowledge is available on later context requests. Running workers keep the snapshot they started with. The Queen can notify affected tasks at defined checkpoints when a finding invalidates an assumption, answers their question, or changes execution. Those notices are logged. A long experiment keeps the method it started with unless the Queen records a cancel or a new attempt.

Priority and scheduling

The worker's importance and confidence numbers are proposals. The Queen assigns its own assessment from the result, the successor rationale, the graph, and the goal.

Confidence is belief in a claim. Importance is the consequence of resolving it. Priority is whether investigating it is the best use of capacity now. A cheap, low-confidence hypothesis can go first if the answer would change the research direction. A high-confidence leftover can wait.

python
class ResearchEvaluation:
    """
    Sketch of an admission score.
    """
    proposed_importance: str
    judged_importance: str
    confidence_assessment: str
    goal_relevance: str
    expected_information_value: str
    estimated_cost: str
    downstream_value: str
    redundancy_assessment: str
    scheduling_priority: str
    rationale: str
    evaluated_against_state_version: str

The first scheduler can be a transparent heuristic. Numeric scales have defined meaning. Priorities get revised when new results change the value of queued work. One priority queue is enough at the start. Rules that stop every lab from repeating variants of one hypothesis family can come later, and they get measured before they become default.

Four compatible labs can run four tasks at once. A finished lab claims the next eligible task as soon as one is free. A task goes to a lab that has the accelerator, dataset, engine, or tool it needs.

Claiming is atomic. A lease names the worker, the task, and an expiry. Heartbeats distinguish live work from abandoned attempts. An expired lease can be retried. A late result from an obsolete attempt is stored, and the newer accepted result stays current.

Result submission is idempotent. Replaying the same package leaves successors, evidence counts, and confirmations unchanged.

Completion, corpus extraction, proposal evaluation, and queue updates are coordinated in logic. These stages can run as replayable durable events. A repeated event leaves the graph unchanged.

Disagreement

When credible results disagree, the corpus shows the disagreement until later work resolves it. The Queen can admit a replication, a methods review, or a boundary-condition node.

A negative finding keeps its conditions. A null result under one setup leaves the technique open under others.

A closed branch stays searchable. Later evidence can reopen the same question under different conditions. Reopening takes a recorded rationale and a fresh evaluation. A familiar hypothesis stays in the graph. Running it again takes that evaluation.

A branch that depends on an invalidated assumption may need to pause. The graph should make those dependencies inspectable so only the affected work is reprioritized.

Stopping

The goal judge checks evidence against the goal contract. Corpus summaries can point at relevant results. Completion still has to trace to the measurements, proofs, artifacts, or other required evidence.

A promising result becomes a candidate for verification. Completion waits on the contract's remaining obligations: required replication, holdout evaluation, constraint checks, and leftover work.

Once completion is verified, the Queen stops new leases and applies the declared cancel-or-drain policy. Late results can still be kept. The completed program stays completed.

Budget exhaustion, operator cancellation, and frontier exhaustion are recorded as separate stop reasons. The frontier counts as exhausted only after workers, proposal evaluations, and dependency unlocks have also gone quiet.

Test

Workers investigate bounded hypotheses. The graph keeps history and evidence. The corpus makes findings reusable. The frontier shows competing claims on capacity. The Queen maintains those structures. The goal judge decides whether the required outcome is established.

A fair comparison is a sequential baseline on a comparable budget, scored on independently verified goal attainment, duplicate work avoided, research cost, time to verified completion, corpus errors, and findings reused across branches.

The experiment asks whether shared knowledge and explicit scheduling help parallel workers reach a common goal faster or cheaper than a sequential run with the same budget.

Backlinks3