Erpnext Docker Fuzzy
fuzzy-program-version
- Author: cuhkfyp
- Repository: https://github.com/cuhkfyp/erpnext-docker-fuzzy
- GitHub stars: 0
- Forks: 0
- License: MIT
- Category: Developer Tools
- Maintenance: Actively Maintained
Install Erpnext Docker Fuzzy
bench get-app https://github.com/cuhkfyp/erpnext-docker-fuzzy
Add the Frappe Gems badge to your README
Maintain Erpnext Docker Fuzzy? Paste this into your README:
[](https://frappegems.com/gems/apps/cuhkfyp/erpnext-docker-fuzzy)
About Erpnext Docker Fuzzy
erpnext-docker-fuzzy
Server-side cross-centre identity-matching tools for CCD Master in
Frappe/ERPNext.
This repository contains two deliberately separate paths:
api_ccd_fuzzy.pyis the current production baseline. It evaluates the formula stored in eachCCD Registration, writes accepted candidates to the Matching Score child table, and stores an escaped HTML audit explanation on each row. Seeapi_ccd_fuzzy.md.api_fuzzy_evaluation.pyandfuzzy_matching/are a recommendation-only shadow pilot. They compare the baseline with deterministic evidence tiers, two safe identifier-conflict policies, a local Splink model, and a hybrid. They never setIs Matched?and never modify the production match table. SeeMATCHING_PILOT.md.api_fuzzy_canary.pyturns only the validated Tiered Evidence High rule into versioned, reversible recommendation records. It applies full-population cluster and source-coverage gates and still never merges CCD records or modifies production match fields.api_fuzzy_review_queue.pycreates a separate optional human-review queue for eligible candidate pairs at or above the selected maximum-F1 Splink cutoff. Every row remains model tierReview; human decisions are stored separately and never turn the probability into automatic High.api_fuzzy_splink_experiment.pyreproduces an approved frozen evaluation for read-only training-size research. It returns sanitized aggregates, makes no database writes, and cannot replace the approved model or queue.
Management POC
The completed, sanitized proof-of-concept package is available in:
POC_REPORT.md— executive case, five-model evaluation, aggregate evidence, controls, limitations, and rollout proposal;POC_DEMO.md— a 12–15 minute presentation and live-demo guide for management;POC_SYNTHETIC_EXAMPLES.md— presentation-safe fictional pair cards showing model tiers versus human decisions; andPOC_RESULTS.json— machine-readable, non-identifying aggregate results.
The deployed follow-up implementation is specified in
IDENTITY_RESOLUTION_WORKFLOW_PLAN.md.
It defines reversible identity groups and memberships, Tiered and human-review
materialization, continuous QC, optional review batches, deliberate rollout
holds, bulk-approval testing, and the next management demo. The code, schema,
and UI are deployed in guarded default-off mode; no live identity links have
been materialized. See
IDENTITY_RESOLUTION_IMPLEMENTATION_STATUS.md
for the verified boundary.
Before 2026-08-19, the ERPNext evaluation approvals and Pilot promotion were
recorded by the project operator, not management. Management reviewed the POC
results and live demonstration on 2026-08-19 and approved the limited follow-up
workflow: Tiered Evidence for reversible safety-gated recommendations, and
Splink above the selected cutoff for optional human-review ordering. This does
not authorize automatic record merging, Is Matched? changes, or a general
Splink probability threshold.
The locked 251,520-record POC snapshot contains 161,112 Production records (64.06%), 89,377 UAT records (35.53%), and 1,031 Fake/test records (0.41%). The POC metrics therefore describe a mixed governed population, not a separately measured Production-only result.
The project-operator-initiated recommendation-canary preview is now Ready.
On the same 251,520-record governed snapshot, 3,528 of 3,961 Tiered High
candidates passed all gates as Proposed; 433 were isolated as one-to-many
source conflicts. No recommendation is Active, and no production match field
or CCD record was changed. See the sanitized aggregate result in
POC_RESULTS.json.
The separate optional Splink queue is also Ready. It excluded all 3,961
Tiered High pairs and 1,097 previously human-used pairs, scored all remaining
816,534 governed candidates, and stored 11,177 at or above the selected
0.938995074 maximum-F1 cutoff. All remain model tier Review; none is an
automatic match, and the queue made no CCD Master change.
A controlled 2026-08-14 shadow experiment reproduced the approved 500 labels on one worker with Splink 4.0.16, DuckDB 1.4.5, and equal 250,000-pair sampling budgets. The 5,000-record control completed with average precision 0.6242, ROC AUC 0.8714, and 73.33% precision in its top 30, but still produced no valid automatic High threshold. The equivalent 20,000-record run exceeded the worker's memory limit, so it produced no comparable accuracy result and is not a candidate model. The approved v1.1 cutoff and 11,177-row queue are unchanged.
Install the pilot
On the managed Docker host, use the checked-in deployment helper. It captures
the complete private db_connector app outside the containers, overlays it
into every Python process container, installs optional packages into a
persistent site-local target, migrates, builds assets, and restarts the Python
services:
/root/erpnext_docker_volume/deploy_db_connector.sh
/root/erpnext_docker_volume/erpnext_restart.sh calls the same deployment
helper after the stack starts. This protects the fuzzy changes when existing
containers restart and restores them if those containers are recreated. It is
not a backup strategy for unrelated private apps or the database. Because the
site also installs the private hksr app, deployment mirrors that app from the
existing backend container and registers its path in the scheduler and workers
before restart; the backend copy remains its source of truth.
For a Python-only revision that changes no DocType, dependency, or asset, use
deploy_db_connector.sh --code-only. It still refreshes the persistent copy
and all Python containers, clears cache, restarts services, and remounts the
backend view.
For a conventional non-Docker installation, install manually:
Install the pinned packages in the same Python environment used by the backend, scheduler, and long-queue workers:
cd /home/frappe/frappe-bench
./env/bin/pip install -r apps/db_connector/db_connector/requirements.txt
bench --site migrate
bench --site execute db_connector.api_fuzzy_evaluation.install_matching_roles
bench --site execute db_connector.api_fuzzy_evaluation.install_default_pilot_policy
Restart the backend and workers after installing dependencies. Splink and DuckDB run locally on CPU; this is not an LLM, needs no Ollama, needs no API key, and sends no client data to an external service.
For a complete transfer to another ERPNext server, do not use this section or
the Docker helper alone. The repository contains the versioned matching and
identity component, while the target also needs the complete private
db_connector app, the app that owns CCD Master/CCD Registration, matching
Frappe/ERPNext versions, and—when moving the existing site—the database, files,
site configuration, encryption key, and every installed app. Follow
ERPNext_SERVER_MIGRATION_RUNBOOK.md.
Safe first run
- Create a
CCD Matching Policywith statusDraftorPilot. - Import source profiles from live
CCD Registration.fieldmatchrows. Only approved identity targets are imported; arbitrary centre-specific fields remain available in CCD but do not silently become matching evidence. - Leave identifier scope as
UnknownorLocalunless governance has proven that the identifier uses one shared organization-wide namespace. Inpilot-1.6, HKID is the approved exception, but it is global evidence only when both values are complete and pass the HKID check-digit validation. Partial, masked, and invalid values remain review-only evidence. - Start a 500-pair run with 100 double-reviewed pairs:
bench --site execute db_connector.api_fuzzy_evaluation.install_evaluation_run \
--kwargs '{"policy_name":"pilot-1.6","sample_size":500,"double_review_count":100}'
- Review the generated
CCD Match Evaluation Pairdocuments asSame,Different, orUnsure. Resolve disagreements through adjudication. Every observedSameautomatically requires a second independent confirmation. - Finalize only after all intended labels are complete:
bench --site execute db_connector.api_fuzzy_evaluation.finalize_evaluation \
--kwargs '{"run_name":""}'
Finalization reports held-out performance, Wilson confidence intervals, and candidate thresholds. Thresholds remain disabled when either the calibration or held-out partition has fewer than 10 confirmed matches. Finalization does not approve or deploy a policy; production activation remains a separate management decision.
When random candidate review yields too few confirmed matches, create a
separate 100-pair positive benchmark with
install_positive_benchmark_run. It uses unseen legacy high-score links only
to discover records for blinded relabeling and reports blocking recall. Because
that cohort is deliberately enriched, its precision is not production
precision and its thresholds are always disabled.
After a representative run has measured the deterministic High tier, validate that tier on a fresh uniform sample of unseen High predictions:
bench --site execute db_connector.api_fuzzy_evaluation.install_high_tier_validation_run \
--kwargs '{"policy_name":"pilot-1.6","sample_size":100}'
All 100 pairs are assigned for two independent reviews. The run reports High precision and its Wilson 95% confidence interval, but does not recalibrate score thresholds or alter production matching. Pilot 1.6 also discards malformed and obvious sequential Hong Kong phone placeholders before blocking or scoring.
Recommendation canary and guarded materialization
After the unchanged policy has both an approved High Tier Validation and an approved Threshold Evaluation, promote it from Draft to Pilot:
bench --site execute db_connector.api_fuzzy_canary.install_promote_policy_to_pilot \
--kwargs '{"policy_name":"pilot-1.6"}'
Create a preview run:
bench --site execute db_connector.api_fuzzy_canary.install_canary_run \
--kwargs '{"policy_name":"pilot-1.6"}'
The preview fails closed if candidate generation is truncated or skips any
oversized block. Only the validated exact-full-name-plus-independent-evidence
High rule may become Proposed; HKID-only High, unvalidated source pairs,
stale records, one-to-many components, and transitive contradictions become
Exception.
Every recommendation form now renders the pair in one protected side-by-side
table. CCD Match Reviewer sees masked identity values and no CCD record keys;
CCD Match Sensitive Reviewer and System Manager see the full permitted
values and may follow the record links. Raw model reason codes remain restricted
to System Managers.
Exception edges are grouped into one CCD Match Component Review per connected
component, so staff decide the complete 3–7-record case rather than switching
between pair tabs. Reviewers choose All Same, Partial Match, All Different,
or Unsure. Partial Match stores a canonical partition of the component.
Two independent matching submissions finalize an agreement; disagreements and
Unsure go to manager adjudication, and a positive adjudication still requires
an independent matching confirmation. These human decisions never rewrite the
model's Exception status or modify CCD Master. When live materialization is
disabled, a final decision remains Pending; when enabled, a finalized All
Same/Partial Match/All Different decision is passed through the shared safety
service to create reversible Groups/Memberships or fingerprint-scoped
Different exclusions.
A deterministic 100-pair sample of passing Proposed recommendations is marked
Selected for QC
Related Developer Tools apps for Frappe & ERPNext
- Frappe — Low code web framework for real world applications, in Python and Javascript
- Frappe Docker — Docker environment for developing, deploying, and running Frappe applications (ERPNext and custom apps) in production and development
- Builder — Craft beautiful websites effortlessly with an intuitive visual builder and publish them instantly
- Bench — CLI to manage Multi-tenant deployments for Frappe apps
- Frappe Ui — A set of components and utilities for rapid UI development
- Press — Full service cloud hosting for the Frappe stack - powers Frappe Cloud
- Gameplan — Open Source Discussions Platform for Remote Teams
- Doppio — A Frappe app (CLI) to magically setup single page applications and Vue/React powered desk pages on your custom Frappe apps.