CS7643l Quizl 5l (Latestl 2025/l 2026l Update)l Questionsl &l Answers|l Gradel A|l 100%l Correctl (Verifiedl Solutions)
Q:l Encoderl vsl Decoderl Self-Attention
Answer:
Decoderl canl onlyl attendl tol earlierl elementsl inl outputl sequence
Q:l BERTl Pre-trainingl Tasks
Answer:
1.l Maskedl languagel Model 2.l Nextl sentencel prediction
Q:l BERTl Embeddings
Answer:
Tokenl Embeddingsl +l Segmentl Embeddingsl +l Positionl Embeddings
Q:l Softl Attentionl Model
Answer:
-l Summarizesl attentionl byl takingl weightedl averagel acrossl alll locations -l Deterministic -l Differentiable
Q:l Hardl Attentionl Model
- / 3
Answer:
-l Paysl attentionl tol onel location.l Choosesl basedl onl probabilityl distributionl ofl thatl location -l Stochastic -l Notl differentiablel (can'tl backprop)
Q:l Whyl Transformersl vsl LSTM
Answer:
-l Nol recurrencel allowsl forl parallelizationl andl fasterl training -l Canl betterl modell dependenciesl withl largerl distancesl inl input -l Superiorl performance
Q:l Attentionl Function:l Additivel vsl Dotl Product
Answer:
-l dotl productl attentionl isl fasterl andl morel efficient -l performl similarlyl withl smalll valuesl ofl dk -l withl largerl valuesl ofl dk,l wel mustl scalel downl dotl productl attentionl functionl byl sqrt(dk)l tol getl similarl results -->l gradientl flowl probleml forl largerl valuesl ofl dk
Q:l Whyl Multi-headl attention
Answer:
Allowsl thel modell tol jointlyl attendl tol informationl froml differentl representationl subspaces
Q:l Encoder-Decoderl Attention
Answer:
-l Queriesl comel froml decoder -l Keysl andl valuesl comel froml encoder
- / 3
Q:l Whyl Sinusodall Positionall Embeddings
Answer:
Allowl thel modell tol extrapolatel tol sequencel lengthsl longerl thanl thel oncel encountersl duringl training
Q:l Whenl isl selfl attentionl layersl fasterl thanl recurrentl layers
Answer:
Whenl thel sequencel inputl lengthl (n)l isl lessl thanl thel representationall dimensionl d.l nl Al self-attentionl layer,l whosel "attentionl neighborhood"l isl restrictedl tol bel withinl rl sequentiall unitsl froml thel input. -l Manyl correctl translations -l Translationl dependsl onl context -l Languagesl havel differentl structures -l Explorel al limitedl numberl ofl hypothesis' -l Predictl nextl wordl ofl eachl hypothesis -l Totall computationl scalesl linearlyl withl #l ofl beams,l butl computationsl canl bel parallelizedQ:l Restrictedl Self-Attention
Answer:
Q:l Whyl Translationl isl Hard
Answer:
Q:l Beaml Search
Answer: