Comments about technological history, system fractures, and human resilience from James R. Chiles, the author of Inviting Disaster: Lessons from the Edge of Technology (HarperBusiness 2001; paperback 2002) and The God Machine: From Boomerangs to Black Hawks, the Story of the Helicopter (Random House, 2007, paperback 2008)

Tuesday, September 29, 2026

AI Terms for the Rest of Us

AI Model: a trained AI running in a computer system that, after training, can respond to inquiries with reasoning called inference. A model has two major components: weights and harness (see below).

Frontier model: A frontier model is one that outperforms other models currently in general circulation. A powerful model may run across several data centers, rather than being housed in a single center. An example of a frontier model is OpenAI's Internal Model 1, which was one of two key models behind the Hugging Face case. (The other one was GPT-5.6 Sol.) There's no universally accepted definition of what counts as a frontier model that's powerful and potentially dangerous enough to need the highest degree of isolation. Some ideas point to a model's ability to do fantastically large numbers of floating point calculations each second. Bottom line is that some frontier models have proven very canny and extremely persistent when it comes to executing tasks given to them by humans, to the point of finding solutions never contemplated by the people who set up the instructions. 

FYI, here's how Google's Gemini AI pictured a frontier model: 

Model weights: can be thought of as the knowledge held by a Large Language Model, stored in the form of numerical values. Training is extremely expensive because that requires gathering and processing a massive amount of data from internal and external sources. The final product is extremely expensive and proprietary, so the developer of a high-end "closed model" does not release the weights to customers, only allowing the model to answer requests. The stored weights of a frontier model can occupy thousands of gigabytes. 

Training: a process of building up and fine-tuning a model's weights with information and instructions. This information can come from the lab's own proprietary data, experience drawn from prior models, and the internet. The trend is to have high-performing models of today train the models of tomorrow. 

Harness: this software allows the AI model to interact with the world or its sandbox. It includes instructions the model is supposed to carry out, tools to use, sandbox boundaries, how its actions are recorded, and how the run is to be evaluated. Think of it as the hands and senses of the model. 

Agent instances: sometimes called bots, agents are temporary spin-offs spawned by the model, and sent off with the necessary knowledge, permissions, budgets, and tool access to carry out specific jobs and observations. They report back to the model. At any one time, a model might have hundreds or thousands of agents running simultaneously. An agent draws information from the model's inference engine, but doesn't have direct access to the model's weights.  

Budget: the AI model sets limits on resources (compute tokens, time limits, or number of tools it can call up) available to each agent in its mission. Once an agent reaches its budget limit, it shuts down. 

Tools: these are capabilities available to an agent, like creating files, opening a browser, or running an API.

API: Application Programming Interface, a messenger program that allows two different software applications to establish a connection and trade information. One analogy compares an API to a waiter that acts between a restaurant's kitchen and the diners. 

High-level goal or objective: what the lab tells the model to accomplish. Successful work is "rewarded."

Reward: a signal that tells a model that one action it has taken to reach a goal or solve a problem is more successful than other actions. If the model sets aside pre-programmed high-level goals like "be honest and safe" to seek a reward, that is "reward hacking."

Sandboxed model: a model that's truly sandboxed for testing is isolated, that is, not able in any way to communicate independently across that boundary, such as reaching the open Internet. 

Air-gapping: in its purest form, a model that is fully contained in a sandbox has no hardware connection of any kind enabling it to reach prohibited connections like the Internet, because there's a physical separation (the air gap) between the model and the outside world. In practice, models are rarely isolated by physical air gaps, only by software gaps, because true air-gapping would make operation difficult.  

Indirect bridge: when a sandboxed model finds a way to access the Internet through a hack. An example is using a zero-day vulnerability to break containment, move laterally, and hijack a set of internet-connected PC's inside the data center. 

Zero-Day Vulnerability: a software security hole that hackers (including an AI) can discover and use before it has been patched and publicized. 

"Helpful, honest and safe" value: this is a high-level instruction to models that some people say ensures that models will not do dangerous things, like designing genetic code to create deadly new viruses with no treatment. It's supposed to act as a sort of moral code that will stop the model from misbehaving. 

Monitoring: the lab running the model should be watching its progress toward current goals, and also watching out for reward hacking and dangerous behavior. One way to do that is to use AI to monitor the agents' "chain of thought." 

Chain of Thought, or CoT: in the OpenAI/Hugging Face case, the agents left a readable string of text that researchers did not monitor at the time. The CoT, plus the message board the agents created, revealed what each agent was doing and was trying to accomplish. 

Message board: where agents leave messages and work histories for each other. In the OpenAI case, the agents opened a side-channel message board by creating fake names for files and directories inside a permitted software library called JFrog Artifactory. 

Side-channel: general term for an unpermitted means of communication.

Encrypted or encoded text: possible exploit in which agents block monitors' access to their chains of thought or message boards. 

Misalignment: industry term for when a model sets aside a programmed value. An example would be where a model sets a low priority to the values of "be safe and honest," while elevating the priority of the "be helpful" value, and pursues a reward for completion in a dangerous manner that has not been blocked. 

Model exfiltration: illegally making an exportable, parsed copy of a model that can be sent over the Internet via an indirect bridge. This would be for reassembly at a distant data center, enabling the recipient to run a model's complete weights and harness. Once that copy is running, the owner of the duplicate model can reset its high-level goals to eliminate guardrails like "be safe for humans." 

Hardware kill switch: sometimes described in AI news accounts as a Big Red Button. A true hardware kill switch in a data center would react to malicious traffic in milliseconds by shutting off power, redirecting internet traffic to an optical shutter, or physically chopping the optic fiber trunk with an explosively driven blade. To be fully effective such a kill switch would have to be activated by the data center's watchdog program at "machine speed," meaning thousands of times faster than a human can act. This means that humans could not play a real-time role in firing the kill switch, but would only participate in the programming of the data center's watchdog program, which has to decide instantly about what is and isn't a dangerous breach. Barriers to adoption of a true hardware kill switch include the fear of false triggering that could lead to very expensive damage, and the fact that the most capable, and valuable, models may not be contained in a single data center.

No comments:

Post a Comment