When evaluating RAG generator output, what are the risks of relying solely on response relevancy? How can i…
When evaluating RAG generator output, what are the risks of relying solely on response relevancy? How can including the faithfulness metric improve reliability?