Zac Hatfield-Dodds

Orcid: 0000-0002-8646-8362

According to our database¹, Zac Hatfield-Dodds authored at least 24 papers between 2019 and 2024.

Collaborative distances:

Dijkstra number² of four.
Erdős number³ of four.

Timeline

Legend:

Book

In proceedings

Article

PhD thesis

Dataset

Other

Links

On csauthors.net:

Bibliography

2024

Tyche: Making Sense of PBT Effectiveness.

[BibT_eX]

[DOI]

Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology, 2024

Towards Understanding Sycophancy in Language Models.

[BibT_eX]

[DOI]

Proceedings of the Twelfth International Conference on Learning Representations, 2024

2023

Specific versus General Principles for Constitutional AI.

[BibT_eX]

[DOI]

CoRR, 2023

Towards Understanding Sycophancy in Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

Measuring Faithfulness in Chain-of-Thought Reasoning.

[BibT_eX]

[DOI]

CoRR, 2023

Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

[BibT_eX]

[DOI]

CoRR, 2023

Towards Measuring the Representation of Subjective Global Opinions in Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

The Capacity for Moral Self-Correction in Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2023

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

Proceedings of the Findings of the Association for Computational Linguistics: ACL 2023, 2023

2022

Discovering Language Model Behaviors with Model-Written Evaluations.

[BibT_eX]

[DOI]

Timothy Telleen-Lawton

CoRR, 2022

Constitutional AI: Harmlessness from AI Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

Measuring Progress on Scalable Oversight for Large Language Models.

[BibT_eX]

[DOI]

CoRR, 2022

In-context Learning and Induction Heads.

[BibT_eX]

[DOI]

CoRR, 2022

Toy Models of Superposition.

[BibT_eX]

[DOI]

CoRR, 2022

Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

[BibT_eX]

[DOI]

CoRR, 2022

Language Models (Mostly) Know What They Know.

[BibT_eX]

[DOI]

CoRR, 2022

Scaling Laws and Interpretability of Learning from Repeated Data.

[BibT_eX]

[DOI]

CoRR, 2022

Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

[BibT_eX]

[DOI]

CoRR, 2022

Predictability and Surprise in Large Generative Models.

[BibT_eX]

[DOI]

CoRR, 2022

Deriving Semantics-Aware Fuzzers from Web API Schemas.

[BibT_eX]

[DOI]

Zac Hatfield-Dodds

Dmitry Dygalo

Proceedings of the 44th IEEE/ACM International Conference on Software Engineering: Companion Proceedings, 2022

Predictability and Surprise in Large Generative Models.

[BibT_eX]

[DOI]

Proceedings of the FAccT '22: 2022 ACM Conference on Fairness, Accountability, and Transparency, Seoul, Republic of Korea, June 21, 2022

2021

A General Language Assistant as a Laboratory for Alignment.

[BibT_eX]

[DOI]

CoRR, 2021

2020

Falsify your Software: validating scientific code with property-based testing.

[BibT_eX]

[DOI]

Zac Hatfield-Dodds

Proceedings of the 19th Python in Science Conference 2020 (SciPy 2020), Virtual Conference, July 6, 2020

2019

Hypothesis: A new approach to property-based testing.

[BibT_eX]

[DOI]

David Maciver

Zac Hatfield-Dodds

J. Open Source Softw., 2019

Zac Hatfield-Dodds

Timeline

Legend:

Links

On csauthors.net:

Bibliography

Loading...