Can Rager

According to our database¹, Can Rager authored at least 11 papers between 2023 and 2024.

Collaborative distances:

Dijkstra number² of four.
Erdős number³ of four.

Timeline

Legend:

Book

In proceedings

Article

PhD thesis

Dataset

Other

Links

On csauthors.net:

Bibliography

2024

Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks.

[BibT_eX]

[DOI]

CoRR, 2024

The Quest for the Right Mediator: A History, Survey, and Theoretical Grounding of Causal Interpretability.

[BibT_eX]

[DOI]

Aruna Sankaranarayanan

CoRR, 2024

Measuring Progress in Dictionary Learning for Language Model Interpretability with Board Game Models.

[BibT_eX]

[DOI]

Claudio Mayrink Verdun

David Bau

Samuel Marks

CoRR, 2024

NNsight and NDIF: Democratizing Access to Foundation Model Internals.

[BibT_eX]

[DOI]

CoRR, 2024

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

[BibT_eX]

[DOI]

CoRR, 2024

2023

Structured World Representations in Maze-Solving Transformers.

[BibT_eX]

[DOI]

Michael I. Ivanitskiy

Cecilia G. Diniz Behn

Katsumi Inoue

Samy Wu Fung

CoRR, 2023

Attribution Patching Outperforms Automated Circuit Discovery.

[BibT_eX]

[DOI]

Aaquib Syed

Can Rager

Arthur Conmy

CoRR, 2023

An Adversarial Example for Direct Logit Attribution: Memory Management in gelu-4l.

[BibT_eX]

[DOI]

CoRR, 2023

A Configurable Library for Generating and Manipulating Maze Datasets.

[BibT_eX]

[DOI]

Michael Igorevich Ivanitskiy

Cecilia G. Diniz Behn

Samy Wu Fung

CoRR, 2023

Safety of self-assembled neuromorphic hardware.

[BibT_eX]

[DOI]

Can Rager

Kyle Webster

CoRR, 2023

Linearly Structured World Representations in Maze-Solving Transformers.

[BibT_eX]

[DOI]

Michael I. Ivanitskiy

Cecilia G. Diniz Behn

Katsumi Inoue

Samy Wu Fung

Proceedings of UniReps: the First Workshop on Unifying Representations in Neural Models, 2023

Can Rager

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...