That may have been why it was mentioned, but within micro applications, that is getting the relationship flipped.
Micro models estimated from agent's optimization problems are more frequently estimated using MLE (John Rust's GMC bus paper being a canonical example, though still true in the current literature e.g. Nevo, or Bajari and Hong).
Micro models estimated from aggregated data frequently lack agent-level optimization, and those are the models more frequently estimated with gmm (examples here would include the hundreds of papers based on Berry, Levinsohn and Pakes).
So, this explanation doesn't seem to hold in micro contexts.
Though, as the previous commenter pointed out, macro is quite a bit different, and your explanation is probably correct there.
Micro models estimated from agent's optimization problems are more frequently estimated using MLE (John Rust's GMC bus paper being a canonical example, though still true in the current literature e.g. Nevo, or Bajari and Hong).
Micro models estimated from aggregated data frequently lack agent-level optimization, and those are the models more frequently estimated with gmm (examples here would include the hundreds of papers based on Berry, Levinsohn and Pakes).
So, this explanation doesn't seem to hold in micro contexts.
Though, as the previous commenter pointed out, macro is quite a bit different, and your explanation is probably correct there.