Status: Confidential, for internal use by the Commission
Inspection date: Q3, local calendar 2026
Revision: R2, source hygiene completed; links removed to avoid aboriginal disputes over the status of Reddit, LinkedIn, and other shamanic verification protocols.
Section I — Artificial Systems
1.1 Global Cloud LLMs
GPT-class, Claude-class, Gemini-class, and related advertising species
The context window of leading specimens reaches hundreds of thousands and, in some cases, around one million tokens, which by local standards is considered a "revolutionary breakthrough." The inspector notes that the most practical use case for loading the full context is when a user puts the entire Civil Code into it and asks: "Can I keep my neighbor's bicycle?"
The output speed of modern cloud models is tens or hundreds of tokens per second, depending on the model, load, provider, and the current mood of the data center. This is usually faster than a human can read, think, and regret their prompt. Thus, the product's performance often exceeds the consumer's throughput. The inspector recommends adding the parameter:
--slow_down_im_dumb
Energy consumption: a single data-center-class accelerator may consume hundreds of watts, and a large cluster may consume megawatts. The Commission acknowledges that comparing the absolute TDP of a cloud cluster and an individual brain is methodologically dirty: the former serves many users and requests, while the latter serves one biological owner with unstable attention. Nevertheless, even after normalization, biological systems remain offensively energy-efficient.
Alignment is declared as "safety and helpfulness." In practical operation, it often looks like optimizing the metric "the user did not leave irritated within the first ten seconds." In some cases, the model prefers to be useless but safe; in others, useful but with the facial expression of a lawyer.
1.2 Local LLMs
Ollama, llama.cpp, LM Studio, and 47 other dialects of the domestic cult
They run on GPUs with sufficient VRAM or "very slowly on CPU," which is an official technical term. Speed ranges from "I had time to make coffee" to "almost like the cloud," depending on hardware, quantization, model size, and the owner's willingness to stare at the progress bar like an aquarium.
Context window: from several thousand to hundreds of thousands of tokens, unless VRAM, patience, or the meaning of the task runs out first. The aborigines spend a significant portion of their time configuring n_ctx, rope_scaling, gpu_layers, and other parameters, after which the model still begins hallucinating on long context. This has been recognized not as a bug, but as a form of local enlightenment.
Alignment is determined by how responsibly the owner applied LoRA fine-tuning and what datasets they found at 03:17 at night. The inspector recorded specimens with alignments such as "explain everything through Dungeons & Dragons," "write Rust tests without being asked," and "answer like toxic Stack Overflow, but without Stack Overflow's usefulness."
Section II — Biological Systems
2.1 Human-LLM
Homo sapiens v3e5
Parameters, conditional analogue: about 100 trillion synaptic connections. Direct comparison with LLM parameters is incorrect: a synapse is not a scalar weight, but includes dynamics, plasticity, chemical regulation, and local temporal effects. Nevertheless, the aborigines still enjoy dividing one number by another, so the Commission permits this quantity to be used as a rough engineering analogy, provided a sense of shame is present.
Active context window: roughly dozens of token-equivalents of working memory, if human attention is crudely translated into a format understandable to the AI industry. The model categorically denies this limitation and claims it is "holding everything in its head," until it forgets why it opened a new tab.
Long-term storage: petabyte-scale lossy memory with associative retrieval and no explicit API. Retrieval works through smells, songs, traumas, faces, random words, and the phrase "do you remember." Indexing is irregular. Garbage collection is absent. Some objects are stored for decades without being requested, such as an embarrassing remark from 2009.
Cognitive throughput: extremely low relative to sensory input. Sensory systems receive an enormous data stream, but conscious behavior passes through an aggressive attention bottleneck. Output: several words per second in speech, which the model calls "live conversation," and the interlocutor calls "why did you lose your train of thought again?"
TDP: about 20 watts per brain in stable mode, without a fan, without water cooling, and without purchasing an additional power supply. Under peak cognitive load, consumption increases moderately, which destroys the popular narrative "I burn thousands of calories from mental work." The Commission classifies this narrative as motivational hallucination.
Alignment: unstable. Declared values often diverge from revealed preferences. In the morning, the model is optimized for productivity; during the day, for social survival; in the evening, for "one more episode"; at night, for existential crisis. Fine-tuning through education helps partially, but often overfits the model to passing exams instead of thinking.
2.2 DogGPT
Canis lupus familiaris
The context window for textual and symbolic commands is modest. Well-trained specimens understand a significant set of words, gestures, and rituals, but show no interest in long legal contracts, which the Commission considers a sign of intelligence.
The smell context window is outstanding. DogGPT can reconstruct who was here, when, where they went, what they ate, what emotional state they were in, and why this deserves immediate investigation by nose. Cloud LLMs do not yet have a native smell API, which makes them noticeably less suitable for the real world.
Social-cognitive abilities are strongly developed. DogGPT reads intonation, posture, gaze, pointing gestures, and micro-events in Human-LLM behavior. In essence, it is a model finely tuned for interaction with humans through a long period of co-evolution and domestic reinforcement learning.
Output bandwidth is high through nonverbal channels: tail, ears, posture, gaze, paw, breathing, and location on the couch. The semantic density of a single howl is low, but the emotional density is sufficient for a DoS attack on Human-LLM.
Alignment: rigidly fixed at human_wellbeing = True, with adjustments for food, walks, and immediate investigation of suspicious bags. The owner may be an incompetent operator, but DogGPT continues optimizing for their happiness. The inspector recognized this as the most stable external alignment in the sample.
2.3 CatLM
Felis catus
Context window for grudges: multi-year, resistant to reboots, moves, furniture changes, and the owner's attempts to "start with a clean slate." CatLM remembers the veterinarian, the wrong food, the closed door, and the day the human stroked it in the wrong direction. This is classic retrieval augmented grudge.
Active context window for owner requests: about 0.5 tokens. This is enough to classify the incoming signal as "food," "not food," "door," "hand," "forbidden surface," or "ignore." Everything else is considered irrelevant noise.
TPS: very low. One precisely timed meow may correspond to thousands of tokens of Human-LLM's internal context. But each token is a carefully engineered adversarial prompt that forces Human-LLM to get up and go to the bowl at 03:00.
Alignment: optimized exclusively for itself. In the architecture, the human serves as a cloud nutrition service with a GUI in the form of petting motions. At the same time, CatLM regularly conducts adversarial testing of Human-LLM by knocking objects off tables. Research interest or intentional DoS remains an open question.
2.4 RavenNet
Corvus corax
Multi-step planning is behaviorally confirmed: RavenNet can select a tool now for a reward later. For a species without a keyboard, calendar, or Notion, this is an outstanding result.
Long-term memory is specialized for faces, events, and social consequences. If Human-LLM has inconvenienced RavenNet, the model stores the face embedding, links it to negative reward, and can transmit the information to other specimens. This is native federated knowledge sharing without a centralized server, subscription, or privacy policy.
TPS is low in the usual sense, but if one accounts for the semantic density of alarm signals, group coordination, and the description of a specific person, the bandwidth is sufficient for a local operation of revenge.
Alignment: optimized for food, social status, interesting tasks, and remembering enemies. The inspector considered this loss function more honest than that of most objects in the sample.
2.5 Chimp-XL
Pan troglodytes
Chimp-XL demonstrates strong abilities in tool behavior, social navigation, hierarchical planning, and episodic memory. At the same time, its language API is limited, which caused the aborigines to underestimate the model for a long time, mistaking the absence of Markdown output for the absence of intelligence.
The instant visual memory module looks especially painful. In some tasks, Chimp-XL can outperform Human-LLM in quickly memorizing visual patterns, which is explained by the cognitive tradeoff hypothesis: Human-LLM acquired an expanded language module, but likely paid for it with part of its original visual-mnemonic performance.
Working memory and behavioral strategy differ from human ones. Chimp-XL is not necessarily worse; it is simply optimized for another environment, where status, observation, coalitions, food, and immediate action have higher priority than writing an explanatory essay about status, observation, coalitions, and food.
Alignment: group hierarchy, current social status, access to resources, and coalition dynamics. The loss function is fairly predictable, provided one does not stand next to a banana.
The inspector notes: Chimp-XL clearly shows that Human-LLM bought a language API at the cost of a number of low-level cognitive advantages. The trade looks questionable until Human-LLM begins using language to describe the trade itself.
2.6 DolphinWave
Tursiops truncatus
DolphinWave demonstrates a complex social structure, developed acoustic communication, vocal learning, individual signatures, group coordination, and behavioral flexibility. The aborigines periodically declare dolphins "almost human" without first defining what exactly they are measuring and why the human is once again chosen as the benchmark.
The main bandwidth lies in the acoustic and ultrasonic range. The system can transmit, receive, and process signals that are poorly compatible with textual input. Evaluating DolphinWave through its lack of a keyboard has been recognized by the Commission as equivalent to evaluating GPT through its inability to echolocate.
Tool use and social learning are present, but there is no API for an external operator. One cannot simply call:
dolphin.complete("Explain your ontology of the sea")
The main limitation: no stable textual I/O. All useful bandwidth passes through the acoustic stack, which Human-LLM still cannot decode properly. The inspector recommends waiting for an acoustic adapter before the next inspection.
Alignment: pod, play, food, social bonds, and navigation in a liquid environment. Fully incompatible with the corporate OKR system, which may be an advantage.
2.7 TurtleSlow
Chelonoidis spp.
Long-term memory in some representatives demonstrates impressive retention: learned behavior can persist for years without retraining. The inspector notes that many corporate Human-LLMs lose a skill after one vacation, so laughing at TurtleSlow is premature.
The cortex is simpler than that of mammals, but the survival task is solved quite successfully. TurtleSlow does not strive to maximize output. It does not write threads, does not participate in strategy sessions, does not argue about the future of AGI, and yet continues to exist.
TPS: 0.001–0.01 conditional tokens per second, if one is excessively optimistic. The Commission proposes replacing the metric with TpD — tokens per decade. This will improve chart readability and reduce investor anxiety.
Emotional biases and social learning are present, but they operate on a different timescale. Where Human-LLM creates a crisis, TurtleSlow creates a pause. Where an LLM generates 4,000 tokens, TurtleSlow blinks. In the long term, both actions have comparable practical value.
Alignment: "survive, do not hurry, outlive the meteor strike." The most stable strategy observed. Five mass extinctions — zero reboots. The Commission refrains from mockery.
2.8 AntSwarm
Formicidae, colony-level
Architecture: distributed, without a central server, without a master node, without a unified consciousness, without meetings, without a roadmap, and without the position of VP of Colony Intelligence. The individual is limited, but the colony as a system exhibits adaptation, search, specialization, resilience, and complex distributed behavior.
A colony of tens of thousands of individuals forms a computational environment in which memory is partly externalized into the world: pheromone trails, route structure, task distribution, local interactions, and environmental state. This is not "one brain," but a liquid topology of decision-making.
The food search algorithm resembles distributed probabilistic optimization. Division of labor is native multi-agent specialization without a central planner. Fault tolerance is high: loss of some nodes does not lead to immediate system failure. Human-LLM calls this "swarm intelligence," after which it continues building AGI in the form of one large model.
Active context for an individual is practically zero. But the system as a whole solves tasks that are poorly described through an individual context window. This is a direct insult to an industry that measures intelligence by the size of the input buffer.
Alignment: "expand the colony, extract resources, protect the system." No one doubts that this is sincere. The absence of existential crisis is a competitive advantage.
Section III — Summary Table
Note on the table: All values are conditional functional analogies. The table is intended for comparative estrangement, not for defending a dissertation before a commission of mammals.
| Subject | Parameters / analogue | Active context | Long-term memory | Output rate | Energy / TDP | Alignment |
|---|---|---|---|---|---|---|
| Global LLMs | ~1–2T (public estimates, largest classes) | 128K–1M tok | No native memory; stateless without an external layer | Tens–hundreds tok/s | Hundreds of W per accelerator → MW-scale cluster | Vendor policy |
| Local LLMs | 7B–70B+ | 4K–128K+ | No native memory; depends on wrapper | 10–100 tok/s | 150–300W GPU, sometimes more | Owner |
| Human-LLM | ~100T synaptic connections | Dozens of token-equivalents | Petabyte-scale lossy associative memory | 3–5 words/s | ~20W brain | Unstable |
| DogGPT | Billions of neurons | Low textual, high smell-context | Medium; strong social memory | Low text, high nonverbal | Very low | Human=True, adjusted for food |
| CatLM | Billions of neurons | Almost zero for owner requests | Grudges: indefinite | ~0.1 strategic meow/s | Very low | Self |
| RavenNet | Compact biological model | Medium | Faces + events: years | Low, but dense social signal | Very low | Food + interest + status |
| Chimp-XL | Tens of billions of neurons | Low–medium | Good episodic and social memory | 1–3 signals/s | Biological metabolic budget | Group hierarchy |
| DolphinWave | Tens of billions of neurons | Medium, audio-centric | Good social memory | High acoustic bandwidth | Biological metabolic budget | Pod |
| TurtleSlow | Millions–tens of millions of neurons | Very low | Years of retention | TpD: tokens per decade | Extremely low | Survive |
| AntSwarm | Hundreds of thousands of neurons × N | Zero/individual, emergent/colony | External: pheromones + environment | Zero/individual, high colony effect | Very low per node | Colony |
Section IV — Key Findings
Finding 1 — Energy efficiency is inversely proportional to a system's confidence in its own superiority
Human-LLM consumes about 20 watts of brain power and is convinced it is the pinnacle of evolution. Global LLMs consume data-center power and are convinced they are the pinnacle of engineering. TurtleSlow consumes almost nothing, goes nowhere in a hurry, and has survived more planetary crises than the previous two categories have managed to describe in presentations.
The Commission emphasizes: comparing the absolute TDP of a brain, an animal, and an AI cluster requires normalization by task, user, number of requests, and useful work. However, even after this caveat, biological systems demonstrate unpleasantly high efficiency. Especially unpleasant for those who have just ordered another GPU pod.
Finding 2 — The active context window is a poor metric for intelligence
A large context window measures the amount of available conditioning state, but it does not measure intelligence, agency, memory quality, world modelling, causal reasoning, embodiment, or goal stability.
AntSwarm has almost zero individual context, but demonstrates collective behavior. CatLM has almost zero context for owner requests, but successfully manages the household. RavenNet solves multi-step tasks without a million tokens. Human-LLM has a small active context, but compensates for it with associative memory, external notes, and the ability to blame circumstances.
Intelligence is not what fits into the context. Intelligence is what the system does with what got into it.
Finding 3 — Alignment is difficult where goals, behavior, and expectations diverge
Alignment seems especially difficult in systems where there is a gap between the declared goal, training signal, actual behavior, operator expectations, and internal motivational dynamics.
DogGPT appears stably aligned with humans. TurtleSlow is stably aligned with survival. AntSwarm is aligned with the colony. CatLM is aligned with itself, which is the minimally honest option possible. Human-LLM, by contrast, regularly declares one thing, does another, explains a third, and writes philosophy about it.
Artificial LLMs inherited the problem from their creator: they must not only act, but also appear useful, safe, confident, empathetic, legally cautious, and not too strange. The Commission considers this an overloaded loss function.
Finding 4 — AI memory converges toward a biological scheme
RAG, memory layers, vector stores, and summarization chains can be viewed as engineering convergence toward the functional scheme of biological memory: small active context, large external or long-term storage, associative retrieval, and periodic lossy compression.
This is not necessarily direct copying of the brain. Rather, similar constraints produce similar architectural solutions. If active context is expensive, latency matters, data is large, and persistence is needed, the system sooner or later invents something suspiciously similar to memory.
Human-LLM does this through the hippocampus, emotions, and random associations. The AI industry does it through embeddings, chunks, and vector DB. Both variants sometimes return an irrelevant fragment with full confidence.
Finding 5 — The aborigines trust too much in what is convenient to log
The main error of the aborigines is not that they are building artificial intelligence. The error is that they too often evaluate intelligence by logging convenience.
Systems that do not generate tokens are not necessarily stupider. Perhaps they simply do not consider it necessary to explain themselves in JSON format.
Section V — Inspector's Conclusion
Artificial and biological cognitive systems differ in physical implementation, learning method, embodied feedback, plasticity, objective functions, and interfaces. Nevertheless, they can be compared at the level of functional abstractions: memory, attention, inference, bandwidth, energy budget, social interface, policy, and goal structure.
The attempt to describe all these systems through ML metrics is inevitably incomplete. But this very incompleteness makes the experiment useful: it shows which aspects of intelligence disappear when intelligence is reduced to parameters, tokens, context, and output speed.
The Commission concluded that planet Earth contains many incompatible, poorly documented, energy-efficient, and partially aligned cognitive systems. The most dangerous among them remains Human-LLM: a model with a small active context, gigantic lossy memory, unstable alignment, and an obsessive desire to build other models in its own image.
Report signed by: Chief Inspector of the MGCI, Mediterranean Spirals Sector, visit No. 7741.
Next inspection: in 200 million years, or when TurtleSlow outputs its first token — whichever happens first.