Explore how transformer models attend to different parts of text through multi-head attention patterns