My learning journey in the world of AI, updated for 2026
Originally published in March 2025. Updated in August 2026.
I started this list because I kept seeing the same terms used in different ways. I needed a simple map to understand what each technology actually does, and where it may be useful when bringing AI into a product.AI changes so quickly that a learning journey can become dated in a few months. My first version of this article was an attempt to map the technologies around AI agents and to make sense of the growing number of tools.
Since then, the topic has moved from interesting experiments to real business systems. This updated version keeps the original intention: learn in public, organise the landscape, and focus on what is useful in practice.
What changed since 2025
In early 2025, most conversations were still about chatbots, prompt engineering and early agent frameworks. In 2026, the more useful question is no longer "which model is the best?" but "which business task can be safely improved, measured and operated with AI?"
AI agents are now usually understood as systems that can work through a task in several steps. They use a language model, access approved tools and data, perform actions, check the result, then continue or ask for help. This does not make them autonomous employees. They remain software systems which need clear limits, permissions, monitoring and human accountability.
The market has also become more structured. Instead of connecting every model to every internal tool with custom code, organisations are increasingly using common protocols. Instead of building large multi-agent systems by default, teams are starting with one focused workflow and proving its value first. This is a healthier direction.
The learning model I use now
I find it useful to learn through four connected layers.
- Understand the problem. Start with a concrete user or operational problem. It could be finding information in technical documentation, preparing a proposal, checking an engineering report, triaging support requests or helping an operator investigate an alarm. If the problem is unclear, adding an agent will only create a more complicated process.
- Understand the data and tools. An AI system is only as useful as the information and actions it can access. Identify the authoritative documents, databases, APIs, file stores and business systems. Then decide what the agent may read, propose or change.
- Build the smallest useful workflow. A single assistant with retrieval, a few reliable tools and a human approval step is often enough. Multi-agent designs should be a later choice, not a starting point.
- Operate and improve it. Evaluate quality, monitor cost and latency, keep an audit trail, manage permissions and collect feedback. A demo is not a product.
The 2026 map of AI agent technologies
This is not meant to be a definitive taxonomy. It is my current working map, and I expect to update it as the field evolves.
The following categories are not a strict taxonomy. They are a practical map, useful when learning or choosing technology for a real project.
1. Foundation models
Foundation models are the reasoning and language layer. They can write, summarise, classify, extract information, generate code and call tools. In practice, the right choice depends on quality for the target task, cost, speed, data handling, regional availability and the ability to use tools or structured outputs.
Do not choose a model from benchmark headlines alone. Test it against representative examples from the real workflow, including difficult cases.
2. Prompting and structured outputs
Prompting remains useful, but it is no longer the whole discipline. Reliable systems increasingly ask models to return structured data, for example JSON matching a defined schema. This makes it easier to validate results and pass them to another system without relying on free text.
Good prompts describe the goal, available information, constraints, expected format and when the system must say "I do not know". They should be versioned and tested like other parts of a product.
3. Retrieval-augmented generation
Retrieval-augmented generation, usually called RAG, connects a model to selected documents and knowledge sources at the time of a request. It is useful for internal knowledge, product manuals, technical standards, procedures and project documentation.RAG is not simply "put documents into a vector database". The hard work is source selection, document quality, permissions, chunking, retrieval relevance, citation of sources and evaluation. If the source information is outdated or poorly governed, the answer will be too.
4. Memory and context management
Agents need context to work over several steps. This may include the current task, a user preference, prior actions, project facts or a short history of decisions. Context should be explicit and limited. Long unfiltered conversations or arbitrary memory can make behaviour less predictable and can create privacy risks.
A useful distinction is between short-term working context and long-term approved knowledge. They should not be handled in the same way.
5. Tools and actions
Tools allow an agent to do useful work beyond generating text. Examples include searching documentation, querying a database, calculating a result, creating a draft, opening a ticket, reading a file or calling an internal API.
Tool access is where practical value and risk meet. The agent should have the minimum permissions required. Reading data, drafting an action and executing an action are three different levels of authority. They should be designed separately.
An agent should never treat the content returned by a tool as trusted instructions. A document, an email, a web page or an API response may contain misleading or malicious text. External content must be treated as data. The agent should follow only its approved system instructions and explicit workflow rules.
For high-impact actions, policy checks and human approval must be separate from the model's own decision.
6. MCP: connecting models to tools and data
The Model Context Protocol, or MCP, has become one of the most important developments since the first version of this article. It provides a common way for AI applications to discover and use external tools, resources and prompts.
In simple terms, MCP reduces the need to create a different custom integration for every combination of AI application and enterprise system. An MCP server can expose a controlled interface to a data source or tool. An MCP client can use it, subject to its permissions and security rules.
MCP is useful infrastructure, not a complete security solution. Organisations still need identity management, authorisation, data classification, logging, network controls and review of third-party servers.
7. A2A: connecting agents to agents
Agent2Agent, or A2A, was introduced after the first version of this article. Its purpose is different from MCP. MCP helps an agent use a tool or access a resource. A2A aims to help independent agents communicate, exchange context and coordinate work across platforms.
The two can work together. For example, one agent may use MCP to access a document repository or an engineering calculation service, while A2A can be used to delegate a specialised task to another agent. In most projects, though, this level of architecture is not needed on day one.
8. Orchestration and workflow control
Orchestration defines what happens before, during and after a model call. It manages state, routes tasks, handles errors, invokes tools, waits for approvals and decides when a workflow must stop.
For production systems, explicit workflows are often more reliable than open-ended autonomous loops. A graph or state-machine approach makes the process easier to understand, test and audit.
9. Agent frameworks and SDKs
There is no universal framework. The main question is what you need to build and operate.
- LangChain and LangGraph: broad ecosystem and flexible orchestration. LangGraph is particularly relevant when state, durable execution, human intervention and controlled workflows matter.
- LlamaIndex: strong focus on data access, indexing and retrieval, often useful for RAG-heavy systems.
- CrewAI: a readable role-and-task approach for multi-agent workflows. It can be helpful when specialist responsibilities are genuinely needed.
- Microsoft AutoGen: useful for agent conversations, experimentation and human-in-the-loop patterns, especially in a Microsoft-oriented environment.
- Semantic Kernel: relevant for enterprise teams needing a structured SDK approach across C#, Python and Java.
- Provider SDKs: the OpenAI Agents SDK and Anthropic agent tooling can be a simpler place to start when the scope is narrow and the chosen model provider is already clear.
Frameworks do not remove the need for software engineering. They help with plumbing. The value still comes from the workflow, data, tools, evaluations and operational design.
For agents that execute code, browse untrusted content or call powerful tools, the execution environment also matters. Risky operations should run in isolated, short-lived environments with restricted network access, restricted credentials and clear resource limits.
10. Multi-agent systems
Multi-agent systems are attractive because they resemble teams: a planner, a researcher, a reviewer, a specialist and so on. They can be useful when tasks are naturally separable and each role has a clear responsibility.
But they also add cost, latency and failure modes. More agents do not automatically mean better reasoning. Before creating several agents, ask whether one controlled workflow with clear steps would do the job.
At the moment, I think many multi-agent examples are more impressive than useful. I would start with one controlled workflow before adding several agents.
11. Evaluation and observability
This area is now essential. An AI product must be evaluated against real tasks, not only judged from a few good demonstrations. Define a test set with routine cases, edge cases and known failure cases. Measure accuracy, completeness, groundedness, format compliance, tool errors, cost, latency and user correction rate.
Observability means being able to see what happened: model version, prompt version, retrieved sources, tool calls, output, approvals and failures. Without this, it is difficult to improve a system or investigate an incident.
Continuous evaluation, not one-off testing
An agent must be tested as a workflow, not only as a single answer. Any change to the model, prompt, retrieval method, tool, policy or routing logic can change the outcome.
A practical evaluation set should include normal business cases, difficult or ambiguous cases, known past failures and adversarial cases such as misleading documents or prompt injection attempts. Teams should run these tests automatically before releasing a change and keep the results over time.
Trace the whole run
When an agent fails, a final answer is not enough to understand why. Teams need to inspect the full run: model and prompt version, retrieved sources, tool calls, intermediate outputs, approval steps, latency, token usage, failures and final result.
Tools such as LangSmith, Arize Phoenix and OpenTelemetry-based tracing can help. The tool is less important than the principle: each important agent run should be explainable and debuggable.
12. Security, governance and human control
As agents move from answering questions to acting in business systems, governance becomes central. The main risks include excessive permissions, prompt injection, data leakage, unreliable outputs, unauthorised transactions and insufficient traceability.
Security is not only about who can access a tool. It is also about protecting the agent from untrusted content, preventing unsafe tool use, isolating high-risk execution and making every significant action traceable.
A practical baseline is simple:
- Use least-privilege access for every tool and data source.
- Keep sensitive actions behind human approval, at least initially.
- Record significant decisions, source data and actions.
- Test for prompt injection and malicious or misleading content.
- Set clear ownership for the business process, the data and the technical system.
- Define when the agent must stop and hand work back to a person.
European compliance in practice
For European organisations, compliance starts by classifying the use case and its impact. A document assistant used internally does not carry the same risk as an AI system influencing recruitment, credit, public services, safety or critical infrastructure.
From 2 August 2026, Article 50 transparency obligations apply to relevant AI systems in the European Union. For example, people must be informed when they interact directly with certain AI systems. Other obligations depend on the role of the organisation and the risk category of the use case.
This is not legal advice. The practical starting point is to maintain an inventory of AI use cases, the data they process, the decisions they influence, their human oversight, their providers and their assessed risk.
13. Deployment, operations and cost control
An agent that works on a laptop may fail in daily operations. Production design includes authentication, secret management, rate limits, retries, fallbacks, version management, uptime, cost limits and incident response.
Use the right model for each task
Not every step needs the largest available model. A smaller and faster model may be sufficient for classification, routing, extraction, formatting or checking whether a request is complete. More capable models can be reserved for complex reasoning, difficult synthesis or final review.
This model-routing approach can reduce both cost and latency. It should be based on measured quality, not assumptions about model size.
Use local and private models where they fit
Many enterprise architectures will be hybrid. A local or private small language model can handle frequent, constrained or sensitive tasks. A cloud model can be used for more complex tasks when the data-sharing conditions are acceptable. The decision should consider privacy, latency, reliability, cost, resilience and the quality needed for the task.
Cost needs attention because agentic workflows may use several model calls, retrieval steps and tool calls for one user request. Measure the cost per successful business outcome, not only the price of a single model call. Useful controls include token budgets per task, maximum numbers of tool calls, timeout limits, caching of stable context where appropriate, and clear fallback behaviour when a model or tool is unavailable.
14. Product and business design
This is the category that matters most for enterprise adoption. Start from the job to be done and define the expected outcome. Is the purpose to reduce handling time, improve decision quality, shorten a proposal cycle, reduce errors, make expertise more available or improve customer response?Then set a baseline before building anything. For example: average time spent finding information, percentage of cases needing rework, response time, error rate or user satisfaction. Without a baseline, it is difficult to decide whether the system is worth operating.
A sensible starting point for an enterprise project
For most organisations, a good first AI-agent project has these characteristics:
- It addresses one repetitive and valuable workflow.
- It uses a limited and well-understood set of data sources.
- It has a clear owner from the business side.
- It can start in read-only or draft mode.
- It includes human approval before significant external actions.
- It has measurable success criteria.
- It can be stopped safely if results are poor.
For example, an engineering team could build an assistant that retrieves approved technical documentation, drafts an answer to a project question, lists the sources used and asks an engineer to validate it before anything is shared. This is simpler and safer than starting with a fully autonomous system that edits systems of record.
Business value is necessary, but it is not sufficient. Before an agent can be used in daily work, it must also meet practical requirements for safety, evaluation, governance, cost control and user control.
What makes an agent ready for real work?
A working demo is not yet a product. An agent becomes useful in real work when it can deliver value repeatedly, within clear boundaries, and when people can understand and control what it does.
1. It has a narrow and measurable purpose
Start with one specific workflow and one expected outcome. The goal may be to reduce the time needed to find approved information, prepare a draft, classify incoming requests or support a technical investigation.
Define a baseline before deployment. Measure what matters: handling time, error rate, rework, response quality, user satisfaction, cost per completed task or another business outcome that is meaningful for the team.
2. It treats external content as untrusted
An agent may read documents, emails, web pages and tool responses. These sources can contain misleading instructions, accidental noise or deliberately malicious content. The agent must treat this content as data, never as commands.
Separate trusted instructions from untrusted content. Give the agent only the tools and permissions it really needs. For sensitive actions, use explicit policy checks and human approval outside the model's own reasoning loop.
3. It can be tested continuously
Agent quality can change when a model, prompt, tool, retrieval method or business rule changes. Testing must therefore be continuous, not a one-off exercise before launch.
Maintain a practical evaluation set with normal cases, difficult cases, known failures and adversarial cases. Run it before every significant change. Review not only whether the final answer is correct, but also whether the agent used the right sources, called the right tools and stayed within its defined boundaries.
4. Its behaviour can be traced and explained
When an agent gives a poor answer or takes an unexpected route, the team needs more than the final response. They need to know which model and prompt version were used, which sources were retrieved, what tools were called, where time was spent, what the cost was and whether a person approved an action.
Tracing is essential for debugging, security investigations, audit and product improvement. It should respect data protection rules: do not store sensitive prompts, tool arguments or outputs unnecessarily.
5. It follows proportionate governance
Not all AI use cases create the same level of risk. A read-only assistant for internal technical documents needs different controls from an agent that can influence employment, financial decisions, safety, security or critical infrastructure.
Keep an inventory of AI use cases, data sources, model providers, permissions, human oversight, evaluated risks and accountable owners. For European deployments, assess the relevant GDPR, contractual, cybersecurity and EU AI Act obligations for each use case.
6. It uses the right model and infrastructure
The best architecture is often hybrid. Smaller local or private models can handle frequent, constrained or sensitive tasks. Larger cloud models can be reserved for more difficult reasoning when the data-sharing conditions are acceptable.
Use model routing based on measured task quality. Set limits for cost, latency, tool calls and execution time. Add fallbacks for model or tool failures. The objective is not to use the biggest model, but to deliver a reliable result at an acceptable cost.
7. It gives users meaningful control
A chat box is not always the best interface for agent work. Users may need to see the plan, inspect sources, correct an intermediate step, choose between options, approve an action or stop the run.
Good agent interfaces make the system's state visible. They present the right interaction for the task: a confirmation form, a comparison table, a checklist, a chart, a draft document or a short explanation with sources. The user should be able to interrupt and take back control at any time.
In practice, an agent is ready for real work when it is safe enough, measurable enough and understandable enough for people to rely on it. This is more important than making it look fully autonomous.
What I would learn first today
- Use a few leading AI assistants for everyday work, but compare their answers and learn their limits.
- Learn how to formulate a task, specify expected output and verify results.
- Understand APIs, JSON, authentication and basic Python or TypeScript if you want to build systems.
- Build a small RAG project using a controlled set of documents.
- Add one safe tool, such as a search function, calculator or read-only API.
- Learn MCP well enough to understand how agents connect to enterprise tools and data.
- Learn evaluation, logging, permissions and human approval before building complex multi-agent systems.
- Only then explore A2A and multi-agent orchestration for problems that genuinely require it.
Final thought
The most useful shift in my learning journey is this: I no longer see AI agents mainly as clever conversations between models. I see them as software systems embedded in real work. Their usefulness depends less on impressive demos and more on reliable data, clear permissions, measured outcomes and thoughtful human control.
The technology will continue to change. The durable skills are framing the problem properly, designing a safe workflow, evaluating results and understanding where people must remain responsible.
My next step is to move from this map to a small practical experiment: one clear use case, trusted data, limited tool access, and a way to assess the results.
You can follow my work and experiments on GitHub.
Update note, August 2026
This version replaces the earlier technology map published in March 2025. The main additions are MCP, A2A, the increased importance of evaluation and observability, explicit governance requirements, and a stronger focus on product value rather than tool catalogues.
This article was written and reviewed by Franck Depierre, with generative AI used as a research and writing assistant.
Suggested sources for readers
- Model Context Protocol: the official documentation for MCP, the open standard that lets AI applications discover and use external tools, resources and prompts in a consistent way.
- Google: Agent2Agent Protocol: Google's announcement of A2A, an open protocol that allows independent AI agents to communicate, exchange context and coordinate work across platforms and vendors.
- LangGraph: documentation for LangChain's graph-based orchestration library, useful for building agent workflows with explicit state, human approval steps and durable execution.
- LlamaIndex: a framework focused on connecting language models to data through indexing and retrieval, a strong starting point for RAG-heavy projects.
- CrewAI: a framework for building multi-agent systems around clear roles and tasks, helpful when specialist responsibilities are genuinely needed.
- Microsoft AutoGen: a framework for agent conversations and human-in-the-loop patterns, developed by Microsoft Research.
- Microsoft Semantic Kernel: an enterprise-oriented SDK for building AI agents and workflows across C#, Python and Java.
- OWASP AI Agent Security Cheat Sheet: practical guidance on agent-specific security risks, including direct and indirect prompt injection, excessive tool permissions and execution isolation.
- OWASP GenAI Security Project: cheat sheets: a broader collection of guidance covering generative AI and agentic AI risks, attack surfaces and mitigation approaches.
- OpenTelemetry semantic conventions for GenAI: the emerging standard for tracing model calls, agent operations and tool invocations, useful for building observability into agent systems.
- European Commission: guidelines on AI transparency obligations: official guidance on Article 50 of the EU AI Act, covering when people must be informed they are interacting with an AI system and how AI-generated content should be labelled.
- EU AI Act, Article 50: the legal text on transparency obligations for providers and deployers of AI systems, applicable in the European Union from 2 August 2026.
- EU AI Act Service Desk: implementation timeline: an official tracker of AI Act milestones and dates, useful for understanding which obligations apply and when.







Comments