rcmnd

Neel Nanda

Runs the mechanistic interpretability team at Google DeepMind; formerly at Anthropic.

8 things

Think Python 2eBook Allen B. Downey

loved

If you prefer a more traditional textbook, Think Python 2e is excellent and also available freely online.

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

fast.aiCourse

liked

fast.ai is a good intro, but a fair bit more effort than is necessary.

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

Deep Learning for Alignment syllabusCourse Jacob Hilton

liked

Jacob Hilton’s Deep learning for Alignment syllabus - this is a lot more content than you strictly need, but is well put together and likely a good use of time to go through at least some of!

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

einopsSoftware

recommended

I highly, highly recommend learning how to use einops (a library to nicely do any reasonable manipulation of a single tensor) and einsum (a built in torch function implementing Einstein Summation notation, to do arbitrary tensor multiplication)

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

Andrej Karpathy's neural networks videoOther

recommended

For an 80/20, focus on Andrej Karpathy’s new video explaining neural nets: https://www.youtube.com/watch?v=VMj-3S1tku0

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

Transformers for Software EngineersOther Nelson Elhage

recommended

Nelson Elhage’s Transformers for Software Engineers (also useful to non software engineers!)

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

The Illustrated TransformerOther

recommended

Check out the illustrated transformer

From Neel Nanda's (co-authored with Jess Smith) 'A Barebones Guide to Mechanistic Interpretability Prerequisites' on his own site.

neelnanda.io ↗·2022-10-24

Every entry is a verbatim quote and a link to the public post it came from. Nothing is paraphrased. If it is not on record somewhere public, it is not here.

The verdict is what they actually said: loved and liked are explicit; recommended means they told others to get it; uses means they only say they use it; read is a book on their public shelf with no verdict; mixed and disliked are kept too.

A dashed affiliate link, sponsored or their own label means the post disclosed a relationship. It is shown, never hidden.

The small bar beside a date shows how dated a recommendation is for its kind of thing.

The bar shows how dated a recommendation is, for its kind of thing. Full and green is recent. It drains, and turns amber and then grey, as the post gets older.

What counts as old depends on the thing. A laptop praised four years ago has been replaced twice; a book loved twenty years ago is probably still loved. The bar is empty after about:

A striped bar means the post carries no date and the date shown is an estimate. No date at all, no bar.