The 5th Brazilian Congress of Computational Fluid Dynamics (V CBCFD) took place on September 2–4, 2026, in Florianópolis, Brazil. ATS was a sponsor, as it has been at every edition since the first one, in Campina Grande, in 2016, as the Brazilian distributor of CFD++ and MIME (Metacomp), Tecplot 360 EX, Ennova and Nexus. We had a booth, and on the last day Guilherme Araujo Lima da Silva, CEO of ATS4i, joined the round table on the use of GPUs and artificial intelligence in CFD.
Summary
- Booth with the tools we distribute in Brazil (CFD++, MIME, Tecplot 360 EX, Ennova and Nexus), staffed by Guilherme Silva, Henrique Lewi (Head of Sales) and Pedro Villela (partner), with conversations with professors, students and prospective clients.
- Round table on September 4, chaired by Rafael Sartim (ArcelorMittal and UFES), with professors Clovis R. Maliska (UFSC) and Marcelo J. S. de Lemos (ITA).
- Our position: GPU computing in CFD is real but still maturing; a quoted speed-up must come with the accelerated fraction of run time, the floating-point precision used and a peer-reviewed source.
- In AI, the most solid use today is surrogate models inside their training domain; outside it, and in certification, the decision stays with the engineer.
The booth
At the booth we presented the tools we distribute in Brazil: CFD++ and MIME (Metacomp), Tecplot 360 EX, Ennova and Nexus and ATS4i’s simulation-based engineering services. We talked with professors, students and prospective clients.

The round table on GPU and AI in CFD
The round table took the morning of September 4. It was chaired by Rafael Sartim (ArcelorMittal and UFES), with professors Clovis R. Maliska (UFSC) and Marcelo J. S. de Lemos (ITA), both also plenary speakers at the congress, and representatives of sponsoring companies. The format was open discussion, without slides.

What we argued at the table
We chose to treat the topic as engineering questions to be answered case by case, rather than a contest over who has the fastest card.
Linear-solver offload or full port
There are two ways to use a GPU in a CFD code. In offload, the code stays on the CPU and only the linear-system solution goes to the GPU (for example, OpenFOAM with NVIDIA’s AmgX library). Implementation cost is low, but matrix assembly, fluxes and the turbulence model remain on the CPU, and data is copied between CPU and GPU every iteration. In a full port, the whole computation stays resident on the GPU. The potential gain is much larger, but development takes years, and each flow regime (RANS, transient, conjugate heat transfer) has to be verified again.
The missing calculation: Amdahl’s law
With offload, the gain in total run time is bounded by the fraction of time spent in the linear solver. With f that fraction and Ssolver the solver speed-up on the GPU:
With a solver 5 times faster on the GPU:
| Fraction in the linear solver, f | Total speed-up | Ceiling with an infinitely fast solver |
|---|---|---|
| 90% | 3.57× | 10× |
| 50% | 1.67× | 2× |
| 18% | 1.17× | 1.22× |
The same card that speeds up the solver 5 times delivers a 17% gain in a case where the linear solver takes 18% of the run time. Without the user’s own value of f, a speed-up figure says nothing about design-cycle time.

How to read a speed-up figure
- Baseline. Some published results compare one GPU with a single CPU core, not with a full node. The difference between the two baselines can approach the core count of the node.
- Type of case. Part of the published gains come from canonical cases (cavity, cylinder, backward-facing step) or from LES and DNS on very large meshes. Industrial RANS, with real geometry and a robust solver running in a design cycle, is a different problem.
- Source. A figure from a peer-reviewed paper with a DOI is not the same as a figure from marketing material.
Numerical precision and hardware in Brazil
Many engineering CFD cases run in double precision (FP64). Some recent GPUs were optimized for AI, in single precision or lower: on NVIDIA’s Blackwell Ultra (B300), FP64 was cut sharply relative to the B200. Other vendors are going the opposite way, such as AMD with the FP64-oriented MI430X, expected in 2027. Buyers in Brazil also face price on request, import lead times, taxes and exchange-rate risk. A CPU workstation can be bought from a local supplier at a fixed price.
Coupled and segregated solvers
Segregated pressure-based solvers solve large scalar systems that map well onto GPUs. Coupled solvers, common in compressible and high-speed flow, solve continuity, momentum, energy and turbulence together, with a dense block per cell. The offload literature we found addresses segregated systems; with a coupled solver, moving the computation to the GPU tends to require rewriting the code, not just swapping the linear solver.
Algorithm before brute force
In combustion projects, ATS4i reduces the kinetic mechanism (fewer species and reactions, keeping accuracy in the range of interest) instead of relying on more hardware. It is an established, validated technique with no hardware investment. GPUs show their largest gains in transient simulations on large meshes: LES, DDES and DNS. They also suit methods built almost entirely on local operations, such as lattice Boltzmann, spectral elements and high-order schemes (flux reconstruction, discontinuous Galerkin). These are the methods behind the native GPU codes with the largest published gains, such as PyFR and NekRS.
AI in CFD
We separated three uses. Surrogate models (POD, kriging, neural networks) trained on CFD results cut prediction cost by orders of magnitude, including for ice accretion, but the gain is in prediction: training data still comes from CFD run many times, and outside the training domain the error can grow without the model flagging it. Digital twins use the same kind of model in closed loop with sensors. LLMs for case setup, meshing or report writing are still research. In certification, AI may help draft, but every number must be checked against the solver output, and the decision stays with the engineer.
Questions before investing in GPUs for CFD
- How much of my real case is in the accelerated fraction?
- Does the validated solver I use on CPUs have a path to a full port, or only to offload?
- Is the quoted gain from a peer-reviewed paper or from marketing material?
- Does the hardware have price, lead time and support in Brazil?
- Does the card have adequate FP64 performance for my solver?
- How long between having a GPU port and having a mature, validated one?
Frequently asked questions
What is CBCFD?
It is a biennial computational fluid dynamics congress, conceived by faculty from UFRJ, UNICAMP, UFCG and IFPB, bringing together universities, research centers and industry. The first edition was held in Campina Grande, in 2016; the fifth, in Florianópolis, in 2026.
How much will a GPU speed up my simulation?
It depends on the fraction of run time that runs on the GPU. With offload, measure the time spent in the linear solver for your case and apply Amdahl’s law: with 50% of the time in the solver, the total gain never exceeds 2×, however fast the card.
Does AI replace simulation?
No. Surrogate models learn from simulation results and work well inside their training domain. To predict outside it, or to support a certification result, you need simulation checked against test data.
Read more
- 5th CBCFD 2026: congress website
- Amdahl, G. M. Validity of the single processor approach to achieving large scale computing capabilities. AFIPS Spring Joint Computer Conference, 1967. DOI: 10.1145/1465482.1465560
- ATS at the 2nd Brazilian Congress of Computational Fluid Dynamics (2018)
- Burner design, combustion analysis and emissions: ATS4i services
- ATS4i publications
Want to discuss whether your CFD case benefits from GPUs? Contact ATS4i.
Cover image: Guilherme Silva, Henrique Lewi and Pedro Villela at the ATS booth. Photo: ATS4i.

