Problem
Symbolic music generation requires models that simultaneously capture short-range melodic patterns and long-range structural coherence. Pure LSTM models lose coherence over longer passages due to vanishing gradients, while pure Transformers produce irregular phrasing despite strong global modeling, yet no prior work had systematically studied this trade-off.