class: center, middle # Understanding the AI coding field ## RSE seminar 11.09.2026 Simo Tuomisto Aalto Scientific Computing --- # Getting throught the terms 1. Model 1. Filters 1. Prompt 1. Scaffolding 1. Harness 1. Tools 1. Model context protocol (MCP) 1. Chatbot 1. Agent --- # Model - Probabilistic next token prediction .middle[.center[]] --- # Model - Served models usually cache calculations involving past tokens (key-value cache) - Only things in the **context window** are used when calculating output tokens - Each model has a model specific **context length** (nowadays up to 1M tokens) .middle[.center[]] --- # Model - Costs of the model usually report prices as 1M tokens for (cheapest to most expensive): - Cached input - Input - Cached writes - Output tokens .middle[.center[]] --- # Filters & routers - Between the client and the LLM there can be filters and routers - Filters can sanitize inputs and outputs to the models - Routers can send inputs and outputs to tools, logging etc. .middle[.center[]] --- # Prompt - Structured input that the model has been trained on - Prompts are nowadays really complex - But main things are still (in decreasing priority): - system - developer - user - assistant - Example: Qwen 3.8
--- # Scaffolding - Rules that the model should follow - A lot of this, especially the system prompt, can be determined by the server - Includes: - System prompts - Tool specifications - User given directions - AGENTS.md - SKILLS.md --- # Scaffolding - Example: How Codex manages the context .middle[.center[]] --- # Harness - Main program loop that: 1. Takes users input 1. Formulates the scaffolding 1. Sends the context to the model 1. Based on the model output either does tool calls or further inference 1. Returns to the user once results have been generated - Switch between user control and harness is often called a **turn** -- - Models can also do **chain-of-thought** reasoning, where they say things like "Let's think this through" - Model reasoning is often constrained by setting a **reasoning budget** ("Let's think step by step and use less than 10 tokens:") -- - Example harnesses include: GitHub Copilot VSCode extension, Codex, Claude Code, pi, oh-my-pi, opencode, Cline, ... --- # Harness Example: How Codex multi-turn loop works .middle[.center[]] --- # Tools - Tools are pre-defined ways for the model to interact with the world - The list of tools can be created on the client side or on the server side - Model does a tool call by adding a pre-defined string to the output - After model has finished generating output, harness will - Examples: web search functionality, document conversion tool --- # Tools Example: How Codex adds tool output .middle[.center[]] --- # Model context protocol (MCP) - Protocol that includes client and server specifications - Allows AI applications to connect to external systems - MCP servers can house multiple tools that the AI applications can then calculating - Example: Letting AI do specific queries against a database --- # Chatbot - Chatbot is a model that is tied to some tools, but it doesn't really have autonomy and it can't do arbitrary things - Example: Model with web search --- # Agent - Agent is harness (muscles), scaffold (personality) and model (intelligence) tied together  --- # Demo: harness capturing --- # Controlling agents - Agents are usually directed using AGENTS.md- and SKILLS.md-files - AGENTS.md specifies the agent personality and conventions it should always follow - SKILLS.md are optional skills that agents can read into context when necessary -- - Best results usually happen when agents are given large amount of autonomy to decide how they want to operate - Using extensive amounts of skills or tools can bloat the context and reduce the effectiveness of the agent --- # Controlling agents - Security with agents is always a problematic thing: - Too much freedom can cause risky side effects - Too little freedom reduces the things one can do with them - Containerizing harnesses helps a bit - Hiding important things behind MCP's helps a bit as well --- # References - [Huggingface's term glossary](https://huggingface.co/blog/agent-glossary#policy) - [Unrolling the Codex agent loop](https://openai.com/index/unrolling-the-codex-agent-loop/) - [Model Context Protocol](https://modelcontextprotocol.io) - [Token-budget-aware LLM reasoning](https://aclanthology.org/2025.findings-acl.1274/) - [AGENTS.md](https://agents.md/) - [Agent Skills](https://agentskills.io)