~/writing
Modern AI Systems

Workflows vs Agents

How workflows differ from agents, and the core building blocks — memory, tools, and planning — behind agentic systems.

Workflows have predetermined code paths and are designed to operate in a certain order, while agents are dynamic and define their own processes and tool usage. They are better suited for well-defined tasks due to their predictability and consistency. Agents are better suited for flexibility and model-driven decision making at scale. Some common workflow patterns include:

1. Prompt chaining

Splitting a task into sequential LLM calls, with each call processing the previous step’s output, with a “gate” check in between each call optionally. These are good for tasks that can be split up cleanly, but trades latency for accuracy.

2. Routing

Routing processes inputs first before directing them to context-specific tasks, allowing you to define specialized flows for more complex tasks.

Routing workflow diagram

3. Parallelization

This is where LLMs work simultaneously on a task. This is either done by running multiple independent subtasks at the same time or running the same task multiple times to check for different outputs.

Parallelization workflow diagram

4. Orchestrator-workers

In this configuration, the orchestrator breaks down tasks into subtasks and delegates them to workers before synthesizing the outputs into a final result. This workflow is basically a hybrid of routing and parallelization but provides more flexibility because it defines AT RUNTIME how many workers it needs, what tasks to delegate, and in what order to execute them based entirely on the unique output. Hence, they are best suited for tasks that cannot be predefined.

5. Evaluator-optimizer

In evaluator-optimizer workflows, one LLM call creates a response while another evaluates that response. If the evaluator (LLM-as-a-judge) or human-in-the-loop decides that the response needs improvement, feedback is given and another response is created. This loop continues until an acceptable response is generated. This workflow is commonly used when there is a success criterion for a specific task, but iteration is needed to meet that criteria.

What are AI agents?

Essentially, agents are a system that has complex reasoning capabilities, memory, and the ability to execute tasks independently. An agent is usually made up of the following key components:

1. Agent core

This is the central coordination module that manages the logic and behavioral characteristics of an agent. It is also where we define:

  • Overall goals and objectives of the agent
  • Tools that the agent has access to
  • Explanation for how to use different planning modules
  • Relevant memory
  • Persona of the agent (not necessary)

2. Memory module

Storage of the agent’s internal logs and interactions with a user. Usually split into two types of memory modules:

  • Short-term memory, which is the actions and thoughts that an agent goes through to answer questions from a user (their train of thought) within a single session
  • Long-term memory, which are the events that happen between the user and the agent. It is a log book that contains a conversation history stretching across weeks/months

3. Tools

These are well-defined workflows that agents can use to execute tasks. Some examples include APIs to search for information over the internet, or a code interpreter that helps to solve complex programming tasks

Function calling

Function calling allows an LLM to interface with external APIs, databases, and tools when a request goes beyond its internal training data (such as asking for real-time weather or pulling live database records). Instead of generating a standard text answer, the model outputs structured instructions indicating which function to run and what parameters to pass.

The standard lifecycle of a tool call is as follows:

  • Context assembly

The system prompt, tool definitions and user message are bundled and sent to the LLM

  • Tool decision

The LLM evaluates the input and decides if a tool call is needed, returning structured arguments if needed

  • Tool execution

Code/API call execution

  • Observation

The output returned by the tool call

  • Response generation

Observation is added back into LLM context, enabling the model to construct the final response or decide on the next action

Tool call lifecycle diagram

Tool definition usually comprises of a name, description as well as parameters.

Best practices for designing/debugging tools

  • Specific descriptions

Clearly explain the tool’s scope, intent, and triggers (e.g., instead of “Search web”, specify “Search web for real-time information or events post-training cutoff”)

  • Reinforce via system prompts

Adding guidance in the main system prompt about when and how to use tools provide additional context which helps the LLM to make better decisions.

System prompt guidance example

  • Informative failure messages

Tools should return informative error messages that help the agent recover or try alternative approaches.

Informative failure message example

4. Planning module

More complex problems require nuanced approaches, usually with a combination of techniques such as task & question decomposition, reflection/critic etc

Task & question decomposition

For example: a question like “What were the 3 takeaways from NVIDIA’s last earnings call?” can be broken down into multiple question topics

  • Which technological shifts were discussed the most?
  • Are there any business headwinds?
  • What were the financial results?

Reflection/Critic

ReAct (Reasoning Acting), chain of thought (CoT), tree of thoughts serve as prompting frameworks that help to improve the reasoning and capabilities of LLMs and refine the execution plan generated by agents.

ReAct combines reasoning with action, usually through a Thought -> Action -> Observation loop. The agent writes down what it needs to do next, calls for a specific tool and then reads the tools’ outputs before updating its plan based on the output.

Reflexion builds on top of ReAct. While ReAct helps an agent execute a single attempt, reflexion allows the agent to reflect on a failed attempt, critique itself, store that lesson in memory and then retry again with better context.

In CoT, the model is instructed to think step-by-step to utilize more test-time computation to decompose hard tasks into smaller sub tasks.

Tree of thoughts is an extension of CoT by exploring multiple reasoning possibilities at each step. It first decomposes the problem into multiple thought steps and generates multiple thoughts per step, creating a tree structure. The search process can be done through BFS or DFS, with each state being evaluated by a classifier or a majority vote.