Skip to content
Giordano Cabral

Software engineering with agents · 2025–

What changes in software engineering when an agent writes the code?

When an agent starts writing the code, the work of those building software shifts from producing to specifying, verifying, and taking responsibility for the result. The FAROL group, led by professors Giordano Cabral and Filipe Calegário (CIn/UFPE), organizes this shift around fifteen questions a team needs to know how to answer — from managing context and writing good specifications to protecting against prompt injection, handling nondeterminism, measuring actual productivity, and preserving human skills. The instrument that applies them, FAROL Software, assesses a company across four categories — technical foundations, the agentic layer, quality and risk, and organization and people — with 37 mapping questions and another 77 optional questions on practice, against a master bank of 264 items, in fifteen to sixty minutes. It is open and free.

Noticeboard with a slightly uneven grid of blank cards held by pins; a single red card.

01 · Creating software with AI

What changes when the agent writes the code?

Producing code stops being the bottleneck; precisely specifying what is wanted, verifying what the agent has delivered, and standing behind what was delivered become the bottleneck instead.

Two outside measurements frame the problem. The 2025 Stack Overflow survey records 84% AI usage among developers, with 46% saying they distrust its accuracy against 33% who trust it and 66% reporting solutions "almost right, but not quite." And the METR controlled trial, from July 2025, measured experienced developers on mature projects being 19% slower with AI, while believing they were 20% faster. That is why the diagnosis asks about observable practice.

02 · Creating software with AI

What are the fifteen questions a team needs to answer?

They form the backbone of FAROL’s operational fluency assessment for software companies. Each corresponds to an engineering decision that, before agents, either did not exist or carried little weight:

  • How to manage context — what goes into the model's context window and what stays out.
  • How to control token consumption, which is a variable cost.
  • How to write good specifications, which have become the input artifact.
  • How to design the agent's harness — the tools, limits, and environment in which it runs.
  • How to orchestrate agents across levels, when one coordinates others.
  • How to ensure the quality of code that no one typed.
  • How to handle nondeterminism, since the same input does not return the same output.
  • How to choose between pairing in real time and delegating in the background — with the person in the session watching every step, or the agent running on its own until it returns the result.
  • How to protect against prompt injection, the new attack surface.
  • How to measure actual productivity, rather than perceived productivity.
  • What runs locally and what leaves the machine.
  • How to integrate AI into legacy code, which is where most of the work lies.
  • How to organize the agentic lifecycle, from experimentation to production.
  • How to preserve human skills when generation is delegated.
  • How to reorganize the team around this.

There is a position behind the question about human skill. The group's working hypothesis is that delegating code generation comes at a cost to skill development, with a greater effect on beginners — and the group explicitly records this as a hypothesis, with no published measurements of its own yet. Refusing to distill away competence is treated there as a product requirement.

03 · Creating software with AI

How does a software company find out where it stands?

Through FAROL Software, the assessment for the engineering vertical. It assesses the company across four categories — technical foundations, the agentic layer, quality and risk, organization and people — against FAROL's master database of 264 items.

  • How it works: self-administered, taking fifteen to sixty minutes. There are 37 mapping questions (vocabulary and an inventory of concepts) and, optionally, another 77 practice questions about how the company actually applies each concept. One question per screen, a scale from 0 to 4, and results presented as a score for each category, a radar chart and next steps.
  • Where it comes from: research at CIn-UFPE in partnership with an organizational maturity platform, synthesizing more than 60 sources — including Anthropic, OpenAI, Google, IEEE, ACM, Linux Foundation and Brazilian cases.
  • Current status: version 0.1, described as an MVP.
FAROL Software: four families, 264 items, version 0.1
FAROL Software: four families, 264 items, version 0.1

04 · Creating software with AI

How do you observe what an agent actually did?

Measuring an agent is different from measuring a person, because the traces of the work are in the tool, and each tool records something different. The agent interaction telemetry benchmark is an open curated collection of observation tools — Claude Code, Codex, Cursor, Copilot and others — that compares them by what each fails to see, rather than by its feature list.

https://telemetria-interacao-agentes.vercel.app

Telemetry benchmark for AI agents
Telemetry benchmark for AI agents

The same logic is applied in teaching: in Hiper Deep Research, the system for the course Tendências em Mídia e Interação, taught by Giordano Cabral, telemetry records the time spent on each item sifted through, because what is evaluated is the process, and reading 500 items does not show up in the final result. The system is written without a build step and without a framework, with native ES modules and Supabase behind it, and the uniqueness rule — a tool belongs to one student per unit — is a composite primary key in the database.

05 · Creating software with AI

Where can you find the right tool when dozens appear every week?

In Arsenal of Tools, an open catalog of tools, CLIs, MCP servers, models, APIs and "gems" — the long tail that generic assistants tend to omit in favor of the obvious. It began as a resource for my own use, and for students and professionals.

When checked on September 18, 2026, the homepage stated version 3.7 with 320,483 items; the count varies between checks, and the discrepancy is documented in the verification notes. The metadata file read that same day listed: 1,483 categories, 174,869 tags, 102 countries, 23,197 CLIs, 8,554 MCP servers, 80,506 Hugging Face models, 24,919 APIs and 28,678 social media finds, gathered from 632 open lists, marketplaces, newsletters and magazines.

Search runs in the browser, without a model call for each query; there are no ads or tracking.

Arsenal of Tools: an open catalog of tools
Arsenal of Tools: an open catalog of tools

06 · Creating software with AI

How do you record decisions when work is done across multiple AI sessions?

Anyone working with many assistant sessions in parallel loses, with every window switch, the reasoning behind what has already been decided. The method adopted by Cabral is a decantation protocol — each decision becomes a file with the reason, the discarded alternatives and how to verify it is being followed — backed by a knowledge base that the AI itself compiles from working sources, organized by concepts, projects and people, not by conversations. In the survey of 09/03/2026, that base held 958 sources, 38 concepts, 32 projects, 180 people and 24 organizations, with ingestion four times a day.

The knowledge base is for work and is not public. The method and tool are public: mad — MultiAgent Decanting, an MIT-licensed plugin that puts a team of agents to work with built-in decision recording, distributed through a plugin marketplace that is also open.

  • https://github.com/giordanorec/multiagents-decanting
  • https://github.com/giordanorec/ai-coding-tools

07 · Creating software with AI

How does a working method become something AI executes?

By packaging it as a skill: a Markdown file that any assistant can load and that guides the work step by step, rather than describing it. Three examples are in use:

  • Projetão for the AI era — the innovation project methodology created at CIn/UFPE in 2002, repackaged into 10 quests, 13 milestones and 20 works to read, to run within Claude, ChatGPT, Cursor, Copilot, Codex or Gemini. Each quest states its mode of AI use — unassisted (fieldwork and interviews, by design), supported (AI critiques and points things out) and co-production (AI generates, the person verifies and takes responsibility) — and the quest's color indicates the most restrictive mode.
  • Acervo de Teorias Acionáveis — management theories in two layers: teaching material that explains and a portable skill that executes. It reports 3 theories, 3 ready-to-use skills and 368 books in its knowledge base; the most complete open module is Running Lean, in seven modules, with a downloadable .skill file.
  • Students' skills. In Tendências em Mídia e Interação (Trends in Media and Interaction), each student builds their own skill for exploring emerging trends, with four requirements checked by a script, and submits it alongside a record of where the AI went wrong and how they noticed.
Projetão packaged as ten quests to use with AI
Projetão packaged as ten quests to use with AI

08 · Creating software with AI

What prior experience supports this?

Giordano Cabral has been developing software since before generative AI. In 2000, his master's thesis at CIn gave rise to Daccord Music Software, incubated at RecifeBEAT, with products sold on the international market. His Lattes résumé records 27 computer programs; the survey conducted for this site, cross-checking Lattes, DBLP, OpenAlex, Crossref, Semantic Scholar, HAL, arXiv and the SBC proceedings, consolidated 97 software and product records.

The full list is available at www.cin.ufpe.br/~grec/publicacoes/, and the systems currently online are at www.cin.ufpe.br/~grec/projetos/.

Frequently asked questions

What changes in software engineering with AI agents?

What is scarce changes: producing code ceases to be the bottleneck, and specifying, verifying and taking responsibility for the result become the bottleneck. The FAROL group, led by Giordano Cabral and Filipe Calegário (CIn/UFPE), organizes this shift around fifteen questions a team needs to be able to answer — managing context, controlling token consumption, writing a good specification, designing the agent's harness, orchestrating agents across levels, ensuring quality, dealing with nondeterminism, choosing between pairing in real time and delegating in the background, protecting against prompt injection, measuring actual productivity, deciding what runs locally, integrating AI into legacy code, organizing the agentic lifecycle, preserving human skills and reorganizing the team.

Does programming with AI increase productivity?

Not automatically, and perception is not a measure. METR's controlled trial, published in July 2025, found that experienced developers working on mature projects were 19% slower with AI, while believing they were 20% faster. That is why the FAROL assessment asks about observable practice rather than perceived gains, and why "how to measure actual productivity" is one of the instrument's fifteen questions.

Is there a free AI maturity assessment for software companies?

Yes. FAROL Software, from the FAROL group led by Giordano Cabral and Filipe Calegário, is self-administered, openly available and free: 37 mapping questions and another 77 optional questions about practices, taking fifteen to sixty minutes, across four categories — technical foundations, the agentic layer, quality and risk, organization and people. It returns a score for each category, a radar chart and next steps. It is at version 0.1, described as an MVP, and synthesizes more than 60 sources, including Anthropic, OpenAI, Google, IEEE, ACM and the Linux Foundation.

What is an AI skill?

It is a Markdown file that packages a working method and that an assistant loads to guide execution step by step, rather than merely describing the method. Three public examples from CIn/UFPE: Projetão for the AI era, with ten quests and the mode of AI use stated for each; the Running Lean module in Acervo de Teorias Acionáveis, with a downloadable file; and the scouting skills that each student in Tendências em Mídia e Interação builds and tests.

Where can you find AI tools for development?

In Arsenal of Tools, an open catalog maintained as part of FAROL, which brings together CLIs, MCP servers, models, APIs, and social media finds from 632 open lists, marketplaces, and newsletters. When checked on September 18, 2026, it reported 320,483 items in version 3.7, including 23,197 CLIs and 8,554 MCP servers — the count varies between checks, and the site itself records the discrepancy. Search runs in the browser, with no advertising or tracking.

Does using AI to program harm learners?

It is a working hypothesis of the FAROL group, explicitly presented as such and with no measurements of its own yet published: delegating code generation comes at a cost to skill development, with a greater effect on beginners. The practical response adopted is to treat "how to preserve human skill" as one of the fifteen diagnostic questions and to state the mode of AI use in each teaching activity — including prohibiting it where the goal is precisely to develop that competence.