Formal Verification and Mechanistic Interpretability?
I am planning to do some experiments on LLM-assisted formal verification with the Lean proof assistant. Thinking of developing methods for transforming natural language intent to formal specifications; and then going into methods for LLM-assisted proof generation given a formal specification and an implementation.
However, on reading some of the research papers on this direction, I find a lot of the work here is plumbing together LLMs and proof assistants, without going deep into either. Has that been you assessment of Neurosymbolic methods too? Or is that a very reductive way to think about this area?
On a separate note, has work in mechanistic interpretability become more feasible? I was wondering if one could find mechanisms to assess how a language model behaves when we introduce errors into inference (via steering vectors etc), especially when it’s doing computation or a proof search. Not sure how tractable or productive such directions would be?