NiriZan¶
Continuous Evaluation Infrastructure for Production AI¶
"Inspection through Measurement" — "Engineering Trust Through Continuous Evaluation"
Why NiriZan?¶
Modern AI systems are probabilistic rather than deterministic. Traditional software testing alone cannot determine whether a retrieval pipeline, language model, or AI agent is performing correctly. NiriZan exists to provide continuous, reproducible evaluation infrastructure that enables teams to measure quality, detect regressions, compare experiments, and build confidence in production AI systems.
NiriZan is an open-source framework providing automated judge-drift attribution, fixed anchor sets with repeatable on-demand rescoring, rigorous statistical gating (Mann-Whitney + Holm-Bonferroni), trust-weighted health scoring, and CI/CD-integrated regression gating in a single, architecturally disciplined Python package.
Installation¶
pip install nirizan
Package: pypi.org/project/nirizan
What NiriZan Does¶
- Automated judge-drift attribution.
AttributionEngineproduces a three-state verdict —NONE,JUDGE_DRIFT, orSYSTEM_DRIFT— distinguishing a quality drop in the system under test from a change in the judge measuring it. - Fixed evaluation anchors, rescored on demand. A versioned
AnchorSetis never edited in place; updating it means creating a newanchor_set_id, so historical comparisons stay meaningful. - Statistically rigorous regression gating. Mann-Whitney U tests with Holm-Bonferroni correction for multiple comparisons, Cohen's d effect sizes, and bootstrap confidence intervals (5,000 resamples) — not a bare threshold on a single score.
- Trust-weighted health scoring.
compute_system_health_scorediscounts the aggregate score when the attribution verdict signals judge unreliability, not just system degradation. - An 8-layer, unidirectional architecture —
instrumentation → orchestrator → metrics → trust → storage → regression → gate → reporting— enforced byimport-linterin CI, not just documented as a diagram.
What is NiriZan?¶
NiriZan is an open-source continuous evaluation infrastructure for production AI systems. It enables engineers and researchers to systematically measure, benchmark, validate, and monitor the quality of:
- Retrieval-Augmented Generation (RAG) pipelines
- AI agents
- Large Language Model (LLM) applications
- Custom AI workflows
Unlike orchestration frameworks that focus on building AI applications, NiriZan focuses on engineering confidence in AI systems:
| Capability | Description |
|---|---|
| Reproducible evaluation pipelines | Consistent, repeatable test runs across environments |
| Benchmark execution | Standardized quality benchmarking for AI systems |
| Regression detection | Automated flagging of quality drops between versions |
| Experiment tracking | Full history of runs, configs, and results |
| Quality reporting | Clear, actionable reports on system performance |
| Deployment-aware validation | Checks tuned to pre-, during-, and post-deployment stages |
Vision¶
The long-term vision of NiriZan is to become the engineering quality layer for production AI, ensuring that every AI application can be continuously measured before, during, and after deployment.
Where to Go Next¶
- User Manual — installation, guides, and API usage
- Architecture — system design and component breakdown
- Literature Review — the research grounding NiriZan's design
- Dependency Graph - NiriZan's dependency graph
- Governance → Data Policy and KPI Definitions
- Standards Mapping — alignment with ISO/IEC/IEEE, UN SDG, ACM CS2023, EUR-ACE, and Washington Accord
- Research Lab — the experiment notebooks behind NiriZan's design
- Contributing, Code of Conduct, Code of Ethics, and Security Policy
- Citing NiriZan? Use the CITATION.cff file in the repo root.
License¶
NiriZan is licensed under the Apache License, Version 2.0. See LICENSE for the full text.
Releases 0.1.0 to 0.3.0 were published under GPL-3.0-or-later and remain available under those terms. Version 0.4.0 and later are licensed under Apache-2.0.
Author: Redwan Rahman — github.com/Red1-Rahman