From 74673b7991dd7b3ed70658b3dd046435c875ff15 Mon Sep 17 00:00:00 2001 From: hectorpatino Date: Fri, 12 Jul 2024 12:41:15 -0500 Subject: [PATCH 1/5] WoE User guide modifications --- docs/images/woe_encoding.png | Bin 0 -> 15995 bytes docs/images/woe_prediction.png | Bin 0 -> 15226 bytes docs/user_guide/encoding/WoEEncoder.rst | 330 +++++++++++++++++++++--- 3 files changed, 291 insertions(+), 39 deletions(-) create mode 100644 docs/images/woe_encoding.png create mode 100644 docs/images/woe_prediction.png diff --git a/docs/images/woe_encoding.png b/docs/images/woe_encoding.png new file mode 100644 index 0000000000000000000000000000000000000000..31ed76c8e142eb948fa5deb48369a700b9d8d62a GIT binary patch literal 15995 zcmd^m2UOHo+Wx3(*u<8^*gz#VEC@sp5D?H9qaz?HQiiI6fK(ZpNYmI76~#dy)WHHs z*P%JIu_ApC1*EGmAkw5W49xtWOOoAu*~`G2+0EhD zkuN=t(p`@^J1NMj$bPl)OItU$(*^u~y4%@z7e0r%^yFl=m z`69afSf!%RlJ?=8RJLqwes5-tjSF{6%V=_Dum1iG!F#4rC^4S~^JMXGhg?@$;jdrp zE}BlEtP2dBNulifa>XpXVB6OVDU@xe&u*kpPEG$9OGBCYISapkvf66e1CjH??`EpB z|7cm2Zo@rz@L=0OshpfAm6Gu_zDwqPh=~(*=gyr@xjk9cH9eWN<_UN16#0%eW!L8{ z^LVvFZ0Yd_A1$+g^8IJ{cca$YnC)k3awpjx(Gj`5InqH=Iwf*D3XIIEukztGi;^605Q zT9ZsuX$#yTUYkEsZ|{X?pilHabNvjwRmt`BNtvEqHIDfW&E9?a{B9e~h%_Ib zw1iLNyjEonUuVOH4OQ{FOzj9InHJ$lw(|S_Y3o7)XKuUnz9HM$qyAX@*0I-Tbn0!m z(NPByu3KJJ_sDebljBHMDt(jf)I@*#!&<)><3+NjKlV1|`o}odwnce1Q!mV2c7Vq2 zmDuH!*I#&fiGq|`?|nbHAj>^v0Z-The+}GunTmRWiUa-MYMlGf~QE z(_@Yw6plU9jJt8;Msd!?ix-!kczv462$nW%zj;ZvH7O z$u6H(iI$>~2508pDAm}kERU#Xe#6ynU+p}FiouO)Qq`BRk2!wLoak~xay-0rpmUJ zF+<&T3RF6lo;Ti=#@+O6Y0zMWn%q3iiQ{qFVGOUq_LPz2+C=>%mhJPK)poDwr7|zr zt{sXsqd(~StW<71zN6s7Ik~nV=7iN=)y~_6?z0wdESe^^G;o2mL0fk^cP}?u-F0%f zwnKr**A}*F3U3w51XAhGcP@(EabYgM#IS&olau52uJUk{NQpCEomPc+pXlNU?Xw!` zrL&iA`RYZSj!BPKhICq<{Cp|xqGy`I{z&g$wihqhkluD}mvvW4cEv6nze2w`bLQ~N zY&3$zBz?Y$XZoG|Xx{S3_~?@xLd`sA!if&e4lKWgjXLFvn7>eDy2EUu*RNS`=JtGN z?16-XiD_wa?!Ap{H#}?T9{h`EUvq4NUVI4o|E#R6jHe4tt`}aFv_PH5u8lr%^ZIQ! zESh3nhhf1$@Bv0u{%CWTq{)KuiqZRUT@CI zpB!n3jq<2Zc**e>RbgSt4{&So9Svq&|7iB&ea82W>`WIqj=ZwGTJ`Flr9>!KQ#4Vf z)Zn0pM|5^bN=lL*?ygJEXz>#$9^>@pbnT(;!3REW*?{8SmzbRFaKJg?;081IYW>*N zv9xq+#atzd#^UT*(<|SS1&LPoR?@H^j@O)AUm7asoM-HcqhW=zrZ#&U*jFd&2XHN$ zygH(&44gA5!PdUKL55Xk!^rq(A9IO^WVUL$oj5ks2%wxD!gh>4PYi>?OWD*EL6gvjHMJ`QkP8+rWeg)+w1(U6kv^*rAt zS>7T-la|!f=&7F^l)8S;@)_AsCl;e#=h2yFq0YhAlD1kf!n?zMclccvqVSTE zlEQ+%t{3sTIi7vd9N}=9PQ99EZ{)RAyA~PzepISXF)V$0n22?sSr3cBjU2U`b?`qc0V2CCEi z6bC!F+%xkuek|}6SU)&5O(#p#H-EoFNx1idq=OrHhMaNz5IKwJCXWWk%;8y!WXo7b zOIg(hCXSW;65`^lXUo0YE@)1?5w9KXs&zgyBKF|?d=}-^I-x+ovMvv=ZX3y8vE%&h z#Ka8>HWg)$=V?Z6vMk)3+TikR*^A1`K=Q}|hF2M*Qn!s*R#cBhmuCx8yq_)eEF0u% za*PG8T=m-f!ZTx!C53^t7XIoj<2Bjgk%o-C;TjW_$)Ojnf}ZS-C-Ze3wieEk%d2ej zJtA$GT85IX+H!a|gH*Bq_zNce)!_@7;fFrC%kI}+NRfD2d1J>5Y&e+)`1H>7HA#3r zL2s^e+}@~gUTLVZm#}m4fjgfUs$AD1Ftlst#HPlQ2JhNlR*V{uXdp|HlvbcRDHyeA z)xXnzLS*nT&2Kn~jL>kO0355_ZkAS@PDDsszJD%`etWomYH>-hw8ZEp$NCKN_&Os3 zHrH2{EWcu{X?p!8QU6tcEB*ja@{TzroEyB6qI7ZYGF2grZjad6aXT8`T> zGNKe~K5REXu+`=5lZp3+oRWBr;e<%vk-AO{=2VVwJa1y4l#K_g72JHIn2?lInv&fd zuIo1t)cNXOtGK|YU4dDA)ofOm-tf&aeLxyRSe%e|=ia8inrF%kDrs&u&Z2?JdDxDBriHC~afM zEB_^?=f};*4y>nkVqzFMHaYtX8V)cHgvy%j!L~ZEQVR+SHhT@URqa*lIcV2X;G1B< z@r$JfNd-5faY#+pD4z3n_2V6#qjw*zwCa*vfsraUWKUDBr<{BDt9zRrUfr!)=GplC zHdml`U(~M5q4l{q`U?4vVipY5H~T3 z&pJ$s&gvvXVXbJ~!ah&^(^CT{Z6ku9B z{glBWjB}OIMpxJFd(Yy4QM~Xc9T3R+!|R7(YMxH2ZtwQ>em-qX6gD=UVtg97As!3E z$MZ)%y|gHP*VW}JowtMN{ZALoGNME*Hk^WeZDGGcf&@lp4a^u zg<@-h5z*`?=o#v1@QjkXwTM;IN?yS6n=CN{s=oWf2|NVwmycCS*SX~3cbM^vQQ?YQ zj~?T8JHSkOLQ0AbUM;ThsQ$0Cu707+cW71%@Z(PnYkX_<7O?ampyNzB}wNZ^S~64qWryuiAjW0Q_j)Os@Tr@%r3FiWt3Aj&%9Ym zb0SK-S>X|LjLS>ECv&DON=+Hq+|D5>Ik~Lp+A0O_k?vt9;0$_S{%AYDgUMYJYRand zb9Z+~H?TKN#bz>I*4D-s6bN@jV$o7`jy^7-hi+pP`^ zM8;@ulN!)oqj9zg8c-H+!^`UGqnBi|igeMGeFlPshFnjQ@t#n66M*>dw-};tq_9-PMfZbip3C;ZrBQn>9OAiffV|ubvbdx@&H8`Z} zWW73aX`|8I7^>jAnX?E-$aFHK(Za|ur?N(yy=|KFyu##3}O|duEaLzK- z%`?R2YJ&~*0s;}3%?9o%^%AF09#|*qpG$7Nq9*HHz_*sY|N6U^FJD?=8!Lu7Ur4Nd zd;*AHkH+U-awTuinQY+jFd(=DQ9Td?qbkxdG z+_=G;Wn8X6C}Ya$KglTnSxNma1^o|x!@gYlBGjm{*~HjQdT3?4PNPi(J(QBTi?6b4M@9K6W zBj^BHIC)dG#kgHipfNl=EP3}NrT%8gc_C zo#_>ha`{75rat{o`Dvo@v~E;$SGNgxIhsIqJ#+)+NKb?Fi8quC+MM8&_S$aiq*#AX ziPiyyLdVB(11foJ_l)o#|#~({(cPK1eqUAO3B*L;Dw_Wk%n_}1*$|I zzdeHe`ny>N^z?$j7u7&UoFiwqRvTLA9OY@sRhR*~c>&J2mTh{rxsH2ok@=6#euEJb z>245?0&MkHNK>V>e+>M1zErYHQ(2fo#6}KUkBr5qXg&Tf?4pY?S#9V1@2=}Vt9Jbl zw3^JEKt=cWFNgYCG?%G)lzn*f7$6}sbHukj%!WSnd{;bZ8xdr9p%w-GpFotWzm2VN zD&V&tG&JmZG)t}wLQV*zE$IuAekT@!yA=wh4U&saO`Abr8^)4`1lt&!A!Db%q6>uU zFeCvcfehJA9;Nz*X2DQycpaJ|zPO+-yTy+$9OW<}JZWHN{ql{#KjB?Xo-O5#Y*&S2 zRksVXKus?%Rgnh*k?r2{J7d1!ZY$+3Kgq#}M#?8d*yMP*l)E?^K%(PAd(CsZuWN4Z zyu`ArP8@u74rOj~k}y6Q8PPSF+gKA{eKy%K zt4_+q;oi=L8;vS&`$Eq=x7TjJzJ6w5M>eK4Db__EODRhJf1+61qP=^MIoR9U76HEm zO^&v3E(c>~Pv|9GY)6Yxfy$r&=?F5!MrdbVgbiXZ)=-HO$6 zaDwjz32~}8A)d*iI8FcS?0I0F_$LRrb6)?qcvog25AxsUQ1g?==4y<+E{jxgwG5qy zramoYgVmbivmeiwKm6#dVyb}oFac+8eefQM8yf&uMl-@~2>dsm@h?T;|A8|3jqLm% z{|PY$i&%u?lXeLSdWnY535<{Q3LytvRvuH*t^3s`K=oY&VFesq50tPl^Isu(|93Yo z&>wxgw(X^!ZaEQq!Mo*mE%&mD_8adEFjc{u7=R88L&Ny0_1RT|?pWZ@yu_r`Jw&!0 zWQu$OCjEzQNgf5S8TE=kGrb$dvl9{$XqKYz@bJv#|1i3Trfm>2YD(qFo@(TuER}q`CDcOeYWgNw zP0L0lhsz=xw^5{JU@19W_yn~ z^jj|O7-(4BI-8Oq+tM)2_#)26XijPssdoJ%!w}7@%3n87<=frC<&7#kCExktc zJ}(ReS0M7x!Of`1awFhDC|;(`K7-|^#~;qb0vE%@i0XV1ue{f(|9w(g9BY8NsAKMy z*QXlYUg+=(l?$#x9@8WelM}j4H;xHU6+`8ag5gh;@ra(m-&tV9lP8N7X%^gRT$AB7 zf3z;{>D9xLrOuX@6)2TCo0TR04{Hc_>Q>lSM5@p-?W&DyYpbgv<~psvy$%ailZY=t zbLA4{wvdO(&>kCVZf9x^;c)MkH)oF|VS{XshA(45`%w+8t${>Q8S=%rwgIfGYq{i$ z?%F225cM?To)7Q4%txAtM|QyXxJhs z&okms!a+J*9ZF@pqMB7c)G`?w!1x2}8+a~kphK}L8mrPQ3AfzW0;0v)r{yaRwwG%$ zx*-p%48PQSkz$hl^(FEkj!Gvy9`2I#C5*1@XsB5TbQ&AzKJ0E*sr>QR54J*vm)XAx zDnQb`hlZ>QjG-z04N75oe`meTS>7>n*%{GTk5&1l}a$AZgCrsYAF|efVV~ z4kB^b2I}q8c(oj%JLI+uBBTHEMBDp{p-@NBJ7YuRZr)r;n5sdlX&9U{2$pSH7UKtG z3&!4tBrm{`465x8RKzFhzTbL<7vKM#LDzKZ<%opZ(OXq6ET!)z zB0mv4z2GL$pDr+L1LSi`P*w2gezkS5OaYzb+g>b1Bjf-gVArzC(GulP4!w)?7xHBc zKujO_*LF9|B!1jhs6};Tq2cayz}cCC`aM|wzaTH zv}M6rh+Y@a5C!*#$^e(lIuU>^F8=6jve@zyUxF%{wXR-7tecY<7>rECaCL7FJa+Tq zy9}$gqH76|gNRUn>xydPT}$-3*woZi7j{;b67_(E1#^6S{Lz)c?fOc|8%dKc9Mkx! zf|tF1B5wvIBBb3TxDRm#d$=9Tyi@X{iIfBWSIzdCFvOZeR22xP+F>#+iX?7S8-%jRIg}TLX^uh2)AT zH5*LLCF&>L7;Qb-EfvJdz_F2Q5j2m=W6@DivR>nz@eyWiPkxy13euFE&V}zBp3hnN z2Y}P*rEyK$-$|nbU;2I_dqSyj%M40Nseoul={ssfFX0#BC$ED0T2y1=OiXIW<6o*m z$)%U9kH-%Ib|-og&wg4RS9Mx3umW5ad2SS%9IT$fUqj`y#T(DsR zS9`rrh=&*)+I{!&Tw*U=s-1&!Q7@lMBfZ;|w18WZleStl=>~9FOMzC2MH9nb!O3g_uUcHbx_GB$%HD~OpT#a|cN)(FN&kE_kGwu#LJqagm z+WVixA^zjI^j{}5{+k|&XtaOO;{;6i8d{xutr_u8tuHL|D8ht{!X`RBNrX6&6$j9U z_6b58T?{Xbj)_Mm54SU(Sx_S+w50h&Yk=v=&R9}zuEYm|ibBSo@{wvHXTL~ooVocK zN>AQ*d}!-H(JCT)GTPQ^@~KEq1Q9v|v}WcQ0gkGk-F(u%bpSz#!Kgl;weY=2T z_%s3YGetXTw>W6)UJL_?83qi(VuOQ&N$kM&<^ElY_L?wXNEfeNa`*QTu-Mi!Zx8-q z16!~n30exVxPr1!)v@qTfXR;4Xb=PWl5CD`RQ6)2;_K`6lkt7OnQ8Kb$OJmSxy^%rj}r&r?!p*>k}eDQf9Vu zi(h_&(`?Y=w{Kgm9W8w%3H1KLuX)-j9W7T)$?L;GrfKH+RY`_v!y^rtU+GrG3}Pbe zI6s@pXoSvq41^idM{SN7KH6N(i4W*;JWr0mOx2^?$fyTpXJ54;?zkk}b@Ckjtby2-3HIkB3SkH^=@4?DK6uq^xU!jokZGIE^K>9+cPH1vB{o(tfa^$L2S>$cB%N6f zW|e(y1MqnLO@sL?@_j&NQVP~(MQ~%U3Hmsih|P?vW)s&0*vY~QrVBB`4r~aPw|jXf z**Mc~l|C6Eic?-{@5<%hWk$H=&JOwR%^1gUSG1vMKv6Nw*qP>Qm(KaoyahV6ck&ha<8?K^U&JfgEoeU0kgA-O=LB{4ue#32en zy;L|4G?7PvTU@MtUtQULLV!>(c zdz6osTn8}p${fMaR*O<~A3wBqU(C-I0^2&PbElLM?rRi7_YxK&_${UZ31-ie;WqlUK;)-_HrD>~)gAtRGn3CdXz+s(F zhCSHM(2}LuFcGKny<>TKnVJ&8iG$gkNf{y?X^gVDBw6>9!*BRo;+z<=a`VBgi>7(B zhA#>rF1;EH7pUA#X1?P7nRe5Z73y8!!ruN z%_PG^5)TNQ5y@+*|3rU)!mJ9WrskQC=9EAeW2}abe}II_@XqbfyvujYQ=baXnK-eJ zVmb9HUf@lYEK^oOIfHLPgrKD!uMQJ9O%iRU;7b8r_%(<$o^r+5bBw{QA0EsgU-tm< zSnAp60wSfg9L%Cx>N!5xx;;r@J8xiaush=-YH|wvMM8dG+5`z3PAN~7Md6u_2sx5;3JIbD;yTxTP27=` zKL#LhU(wG8@cNA7=dk`%IDC~$!M#f{3J3qzB7Vz}QW`Y+pUme!oaYxvxHy;f86jP9*uiY7V~i0F^bLwg1{?;j1aw;M0+ue-}4%1gwfQ5-S}3s+6C ztc!vAFklX%{0&Vrb6+qj`Z)VEdw` zuRUZur?pkomYU)1cz1svT!__=_W6zGl9}v%MAlQgNQxZ!i0jpL} zpz*DuiNSKe#h6AhYceiM)2L+oh%RSuh7_LY`kq+#UQ6FD?2bnA|JxN#eDnwB!+u6b z6eF>pf?9Pme9%6qv`*xWGlwzvdb^dIho=wJ}BecT{ z0bO}GG^>kAfQ>}=$haFH>B+5eRC9lSUBh?WA$jVXQuWU!6OSDrCSaO06*hesAS6*< zRNmjYRg=?$fj~#>2~MpBe<)rvN;K#%{cAfs+hoo(O3WOSsQ^r{qLohMmQ4^;id6+rfBI(hiLfBf z6MgC~K#R!-ndp!y4O2n} zZAT4+vrxUgDo3Mykpd3H!dSg|bWW$<5OfQoq;utDM#!=xglNk>?mhJnwenQ3nr9+4o|;3dlZu z)SAafM&XQ0^Td+Yx5-_f)Chjk#s4NlNMM`wWwGTTaKX7>7=I71_3~gl@qP=v25l~v zJ|>@@GL!!6D9nGg|6jc$8ZepQk**c%J=cHnj)(hrJpO+tpT87U^ zdH^|4q!s%RybFk_*ITvj7Wg1Z_-Z51Kc!57$Q423HtuDuD*^{+AtD8lLbeXb(iK%t zf@^Y~w2VNEOA?j_w6!GpIh?BNxwn+d`vd5)H`o1S6LFRZl}U7s(ot99Ki&;P>H- zq!ol;S>Lo|3Vgm%f*2+}>seTJ+$>j?SC_2r6)5fd45-qsJ>*jk5Ic_HnX4sE)fd~A*W zA9SzhF8yS)&)A1qGcPvR9o z)z=^X-q%N1*{V>v5<;Dc5}&0XiZ z=8Mjsa}p6gy|~%%OFB~CphhV}+H+dQ91`On#7zm_wOz|`xF3j)Lk_b-@TLSpub!b{ zC`tWuP7-&HTs5hBHA;fgxWf|+&avk4`8oCuEwFtCkON9{;WxAhJ&338Gn(7t=pf*8 z{iD$scwB#>tH${At4?PXTF;z-*+R$FlQ<+w(YgGVW7ntB+sL5q)!?_$Lw7wC7Z;y0 zgmB-Z4CmCmm%gs@_YxYx35i)dif%q7SE={|p`1uC8FD3yCo484ITtRCDe^$H0v|7O z_%rNlWr-@wutId3-|NYVARmFqzc`o?~mfM=K-UFCf9q_(CXDRIy@N&N5cs-gLeP;;Sp>>VU~W9 z0gp6}1%{^A@UtB5JzaP-`P2F=qWj}y>>X&Xjw7*|T61qlf+>7Nf)*}Gshj%od36ZY zkh6;pq>`lL^c?>yRazP%Q%c$p#^{=)WzIMS_UxcHJ0SP<5L;gHm;~f05(<~s|3-@p zpbQPNMB8MHrqU;{?dCl@#zYwq6x-n_=!k@qFH)l{ z6jO?E4eNfrTeCmDFC8h*kc5>N4Up5Vq1o^t8I&TCl@84zICr)2wqKjs^Y<$v(~T?( z$%B-bm`WUaJImM<`EFA2%KY~T0>J|66!w)Dfp-m{%eMWnHhMQ>s&yc+GJshblKqv^ zD?}6{$R{Q~1I`2k=cNs){}PC~bm%?FBZkivNf{H1gBXPbehs|6B-c(7{v_`Jy)NMR zLZ5x`*__D5xAxqYNhk9YZ_W|mOw2V`=*l6vn?IrB&vueWC8;(t{h-7uDk}1E`)WU? zfe}(lsqFS}Q~$|fa@D0^C{~U%R1$(@Yc+K92UqUoMq=M$G zWb!s3c~mQ^uFeA@b%y3HpQ_X*HhAQWvuP{Rae$|uBSLi z{fu2G+t`tqt#x)uDdg$^8I!sgx0t|`lt=)W3>qnna^;ptIp0n#mABuoqC>)0q=~Ro zqz!06!~KzXO}K{Ri|dqR(iCdU(-=FXA?&|ugX;bF(y^e5XcfLrM5s(4|h<&trDryse4Am^uQ3Xn@}ZZ3SnE44r`74D_zm93vt^+YL#L13*`Z_uNF*Hp_ulv8DAg>F7x*4G^+FNm_8KW@pj#m~tAMe$)^5 zLfXYN$F?IAS8dgVsS{=~`QDJ&HLhrbO0-dk9Ghv)h|msV`jIP|PW|E2Y}L00&D<6- zbxR7Aqx8|aSp>dHb&;mj0Zd*_?&px5A%#i$=T-O%;XOL4v=f zf9!?dcq;JShX0AVUDtL~bF^^tFm^FVDHyvs**Ut|S(}`?WA5T=?dTxDE5^&ubIQui z&B;}qkI(+^5AZs=Sn>((zF7)4+2?dq#}$R5HAepKNJpnxqfp{gm}^(mJ)`G`eD&4M zx3-toP%cUQr?^h-8&2kYm62ufD&;u0O+;>K&V>@QRK3TkuTsLwbnNaVWv~s|ED$uS zQnDQI?LDvLi*ascXR6OOFy`0(xw~;bdW3ayPSNIK%iE99ealOtA@*%zePVIek7?s( zE8E<80*p`WM4=3uo4R+PPlxa)t@QF*~O|~TB{6N0(gUrlKL8sBnVPRqT_6SL@ z!rokitUB_%<3OQ#&@)za?%}gH`a}9~xarg=2UT|Z4JEgg` z8#XoDTcpaLlTq%Q>(1A2+_)8bNkCxr^xt~AZ3ePh=(J1!w+^JKSckI|v*H6S{ zsD&nf|L%n8$~CA`*U(7$@ZlOs%%IXu#m+8U3U6s?nPuCTOT5O1fj^atwWO5>dZTpN z5SEfBRVo~$@um>&bLSenVIk?iii^t=wjaPXmozsw_te_w?^obMI zwl{}duYXz-E}8#X@?&|~yhL}NWUJN_0&5`7f2*R@VMuq1RFw!1Uu3}9dz&?Ui*`4! zh2QIQu{5v~&}|I1j}rPv4+Y(qA^5RI`L&2?X2}$8vu9*v=$1Le6c!dTVfhjh6V-Ke zGQNH_2_)3!=PRqIs1RMHe}1D2oQ;Xe@h3(PoRZmzl22&Nuc#KQ=ci6dO||LC)EFqX zemPL)xad@E*&b`tmHb%P;fvKscvMu@S=ndjPoKUL`s(@f`0dRpw7A!gWODRDM&XY( zC0x?dwY>#`gwHIfy90%vo-4TfqgwH$F!X(&!?;&N{5NG_;)xfD9UaOrbew{M>ZXlP zRdjVD-`Cf}>kApb-_6d(HciONd(C!`NmO}hX{mC#O+v_VSn}A>qti=GtbOZC>N~RCW^UEiXJ-vH@m%G82?RQcC%tWMkUJJZs)FDVJ}`J zJbn7~Kx}uZy!BP&u6%xO{ldn%WZI?>Wd1r;O+2I%(wQdSYI0TE3o9d|i#<)VxuVWqJ* zpm@Bw8tiQ0n&NBP@J0$g{EYJ9r?A;aC+rH-ZyXPYi`>SMf2Kq2Nxd!5W|v#+nEmU? zoC|*4M@0JXOD%HeFvwHqpFEDIrneZc>gesBhUa+OtO7TczvtiswM#$PM}7Qi;GJMb zQ)?^ABc0%VhG#^+_HNLzFe%jAqbG`xYkE#dXI|7lQRMicd7>+r7Iil~33=xG3diyK z>CW~x)hSYx5wKt>`u+9$IHo({({HcOx4fx2!1q1@85;fUUvJ>*Z0}o%{8iHSuDT8FBtsg+o?GLXbxawzo@OVrU(ie`n|W zhQTj~)*n#O_Kv=Lz~|{RSZaR@gHeH{$c`Cr4DOcWt4UUp{&}tJi%XDVjCe|)p|=9J zRxa+K$d^+X-AXsukzb2_Np}pOj+Q)o2y=|$c=hU)sNJz+$J{I1g!=T8glv1K`wL7j zQOmIUttSxY2h;BDqPi7&XMP|hGn0$`+_`~4x0$Z%1qB74+uH6B3@NkLPt6)Gf_d`z zv5J*dy04s)Ql6S*(D4hZK0ej3r_cKP`yXTwym|8E$;z#bS#}KU!XImETH4xgWU4-X zya}63)SVAms;Q&74x5{sJ~}y?q^-Mm?}{5Zb|zjQsP(Pt&DKeRuEE94Ee9K0%gwD= zMre@+X@JqUOLUn-eGN9h*VAHGI%P|2dboz%x)nl`x$~-3y@_+3MeWjAP`0mMzn)Q8 zsM^rbK-1%EY@7f^D!P)DGPL=hN>2V17HL$40uOKbb6qsm*vQBTZ)drHE3#A)vhKRZ zDB@7e+TPmQ^Wv;5RH0EEnX-yTHxk9$N-kWu&;!FX7q|T5N3mr)%frAxZed|f^2TcS zD^9heM~^1Mv)e$o)2;QXi1pv{!l=SPH9kF=R9q|wA8xiOv3dXJKBy2*Hs@-u!P<#8 zocmItwrekSzQ3ZJ+)K8UsiFsNLy%sYJIL^@J#cw5x;C{y|7r!SH6gP@za z|3&NS7qAs|6DXT)Hi#fNyz#~iDJ)UOFIqo|^pT5wO?y1I0;?8lF)S3w5C(7%3y8CBHWGq}G>Eq&6v!f7M&mx1O~Q7I8tL zl)2TBIuuW-im=D+c-^e3x_Y}A9KWGZ@tTxNl})vw4rx^*rIqq}jTC$p5ad`RpnZk* z@0y`}HpoWbGa~r!&qa+5lvlMnnerceN-1beVKtv>D$ZYJt)oxf+_mEeZ82wUpLMX; zT)Bz3W`BDIPucVYzH^{DwIY*?hk)PqR=A1IoXp9bpNf_oxE7RJ(m5O=bF@GEM%~uD zMH(q2&Yd~fl$1-Bh9)DdeMRyfQmXO1b>caZ3^N_{SFkH<%<4Ld z!8mQaesuiQEj^{G^Jm&WepIfyGk+B3SlV}E8dl1Z+s)La#kt!DYi9K>XqOqyII^fK zl&nn8hi7iC`1=c9vNTAXXS}+)E@w20_n=Kr&wo7m)Y=klsCHIiTZT$QW5a4Ik{tQt z{b74g34KS&?txt+f27wmSH|=zcpQZXxkp19$G)EVbS$u#n_HVM(5>sk>6l^*Yq^uc zw0riK<$wD6qgUHY5?{C4t9J`(VM~NSiT9d=`)scQ04Subr^yIBOw(lPdOLC~!(htA z0P9u3qVIH#Bjl&;cHILuj5)jX#%SKy#(1_VVYo~^CU^6SbvtWkxWt37qO1%y>$TU1 znDO-Kvd39_Q?xeYhedqy(2_n^ZNp2{HO$A!A-jezoVMt1IX z&hg7h88&3fs3+}oU%4G$*c*_M*)?8w>7Mw~y3)dOJ36nxw8>^=y7RBfA1Q2~*=OWG z)SVEas`Dwr2z%OCC^VV8?T|^}VqGoI#rY}PrR~42ubL5U(8`7}TChvKWAK@|a@Biv z+c2$et!GTX>B+uw?PO_p`8c~`LQ%+7CSqV|x~hJAH`f*#-D*^1q*6HL-pR z%4)NvVWG(q-TPO!o*c*WRMgOHa@_RC&OT2Z)QWbMW(X-VCFa@%hsmDZCJs7ksHr7E z2@fQ{c=3Yk{CUx-%_mRxhaoM+DXt-;v~w))zM%alW2L~IxG2reuu8`^XQ>#5tjs5_ z{(Cky%V%rqiaWNwx%H{gj+x;W28l^jyIu)qrr%4tZuts84~+c`7cY25tZ>(Ji+v;SAglC%L3a2gqn zxNhCzDMbk@SX)~Y+u7OK*?j0v2E1D8XCEk)R&U=V-%T`O_K$F{{879a z*}sDCy=~yp&55z}n~9N}%qx#I4^3V-G{?FcCqJoMRnu-3_~h%M`6w~~cb-`IBY=GM zEnok#N`B+P@)qNf2g}Q0{-=~3EsMJz*KchQ|BM^ZC#)BFV*>Zje=)QFs61v75F|i2 z$T3L6BB>4Fq;_Sr)IlwS!gcAG(( zu{ChW$_+x6Epw18w>LnA)j7Nx(rp)u@OA^`&Phj(9BC6n10&PS>M?#`dmKG*BWn0y zAMR=W{{5xM`hKFgk?}Q90F7(S7G@tZOb&-4CdEpHl#7)w-?ST6dDLK2g(e%Vk_-#N z!)%BD)Mnm1qbwm3Xoc@qVZip4{h+n_AgB0_o2)#URbO$1M(?LJtMB$u1UdA&ztNF| zG1FR&k*FufhH|g^zj3*GF?#J-ZR=;dKcTRh zS4~wl6sRmNO^G!ycMCRaMrP(`V?(`gUW22IjI%4%!XfRkPy*F%-FkvuC8T8KR^B>r zoJSd!%Z(70`I!!R;~~Zi>=?BhcU=0nM>o$htXBz)zg%k*KgSxBz|mN$rt$E4&4zy- zZY2rf87X4TwAP)t6ACQYh@DXbuhs|8q{xePIvMSDO47(HRegCmY9Yo~AIg;Yig|qM z(C`%w8Y)k5LH*r+14)vulk%;RLd_Yr(UM-ODJdx%-EmvV6|Pf+nY6fG0IUcQ^<5te zB=qIQ?ccxu+u~wgem;L-DSc!g(N!{b&VQ|HWCmDwC(wCyW##yVp(QHx$+PvwIGnuUfGPkG|bO znyRjroW~c*j*X2y2@dW7=(kceLRN$}QL~r&MEvS(pY8lW$w0MN5!?S_p`t=`MrRPj z^WI=+y`;o-O7Y+c{?zwn&K~)$i0K^q zB@_g6+4rK|2)GQacl3r`Zu!guZ(MSOe@vd5hGvHUxf zUN$=f86vp+Z-z^n<2&~YoqDzW+jWvxLMV7mz+?Q^mrs6)VcKBZTlD`lv&JcZ&^nzu`y$p2GZe8Wd`{nO{FTc;r#qc=vDDL~u+a>Rak% zxar&G2(SArzrUZCurvj@T87qen-U$UQ}6x$eG!)9YLuaKs-JwCoy~KQSzNudvlB!q z-5Tubo_2)qn-^#o-+s3Qd!6OVQu>K7am<%h3}tH#6Iu z^UBi9=>np5jG5WDt>D`?GvB2KdPJ4$&CyGPF9~K8X`9hTxc8f@G9Kf!zl!u0-t`%? zMCv}ebJGuKr8^WvhQ2xRaB#@Z_T}}Z#rb6#)?hKA$1mt6pOy)j##_coEwWP4Qy@N- z$m)3bzewZ$CX1q(Hi<U2oBnGGi(L~^TqATfzmG*DVgEST`#z1B0Bz#4|ADV(w_lee~tq z@grVSF}LL9-TL8mlO-0b)skdGUbBH#lwMx04Zs?y4eXdGVf)Zl3$-|DU!arlmUpqr zAQkMQqUlFZO-xwBXOPzbxezphF(<#5=M;_}JC@QGBOwI+Fs~I4@jKkixEXK&7?u4L-)LnJyR3 zTe;N8C}j5W(CYl47GK?Fp_=neg!a6&nl2k&T^unApbM@A?(IqhwrM-@Ijn7HT@s-_ z-(8T8xu!SoTtnNB-0b1P=jjxs_R?2Ryw)QyQKL|+6DE?vJeqBgN2>bOh zhubIi@bjXg+JQ98p==6q(gp|~7~HiR_4dM-mENkxk27Z5-(D)&0MU&Uag5z0ljli! znM9op23lq3Sab{vQ*~AmLSRJUWa*^{L?dhVzZCw8Em zm`#Z#9_>6w|g;8t8*T&gN6FI-$)%!=@C@Ozx>Qxxj$9z`u8A|!$MfU9t|B-O&A zvjst=a-nke3j-4j$T)#bLO`J-=<#E*#jj#8O;4UZyCbIzBOTQsT2#?hTi{a3xE)~` z+xh1{x?<=?q*_>tgnLvh9iXbr_Kw2HG1Arjk)|2u3wYMNV6%*rn4K-}yRq^jZX{x| zr>6%r=5CNbbHH%ulZyw+?Eq_uo=01`-58cl`-!4$FX(Km%pg&&Hh%~i?(tG##&E$R z0-vB>I;)4EdKq4&COS4=OH(ry6u9fxuBGb~S=bC$SJdwXZST_ux|L`}A}x9|ALR99 zcDdK*0W9X#Ze&`t$L35=TU{`!t(G={Rpx$9k*Me~o@Y>{*&NQN1+tkCcE!p7CJ2^F z(m<(w7VM_;Vqz82Uw5NW?%&+ZjnvhXX_&=tiVH}?k8BviU|AkW++P@(QK$uzuUcQk zW9B5>rsJXUV^jgztpFvG6IWY`^Subml|HJ7d%B`fNK=u6)R`n}yEt4Enw+7Q#&!0r z%)NW}BEcyEU*#(iwF7n6Hkzd<4t_z7qobn(m&3@qFV_Hhtla7!x61ZXq3#ZTkMG4)v*g>h(qzZ_nR0&PU4 z)i;B2Z*Qxu(++;Lcp5E7rxdsY7v#7wJ3n*iE>=xVhqq=UB%tCX-}@a%U-P^l+Z{3t z57zh)>9B!&4H?lOzR3jAXn^>}b38~h&(I#N34)m{i`O-%#Rwrg%FLXD-W;P9_uDM5 z7kcvK$*bsSo~#}qWBkD<1#SSYRP=WbMLIqMc27-V;U$FkgDg13B;mVJd}nbOAw#xn z3q#noA&)T7{q-Cw8pf*sM^9U9L!|1}+uujuF-UrT_X>F!5MWauaFB*{oX;>D!PBzO zm=K=^_C>KD8SA?-lSb&t0N43?O~em_d0ellqB8dSJBX|7K@BGU?egG6y?_6H?y;eGkxNcU zNXVVJ{xMuhj3RNddihysXec87wM&y)w%CurBe_~Xj?*V;aYFl1jR`c{M2hbf|S=;@|Z{bVD=av>6h?|M$ z6t+lqq>?_vtf49o32!kJiXjY;bBwtA=UVXXT2p2~D={!2AhW!@pZOvbpgXy>sYE-ytR9=ZS_60CPF86jZRl(+LJ`8n{!_Gv+KR7SGFhbz7CNvJl5`Zq80>r9&oC+PAVO zd>1O-Sn&IPP`^@W8k`LcV~KP4BrxTmEbEpHW3!Nj11l*bHMLQihLwed7g!|1EIZ<_ zFi70V22&8mAq7fDFoV$NjEHqr=1?*id*)Cp$Ad)(D>L1QE`99Kq4+=+>`l6J@|0Kc zOo38o^Y-ozVmq+x(BxOIjsiuJv+d1(!REd=Jd`jL$m*BAMJCCGYG!LE!0m@T$AkYe zfnu<9^E!$g-a!qC2p3jMmG1%@~kC$DwF7)fU5(xeHZpCA5;{K~IdP`PTR`-Q3I2Gy}%KI2Oae$idDgtU5CN9ocO=n&0Zs;x(cRw@~$c;>@n zIxbF5=YDwJhLm1YBJ!1hU(*4l-VAl9m^dD2?s~BPd4q`xd^afxias-gv-`pmT@Dhy z$Z~erd$9**O*?K0dJo#OZ7V#+_n1+gADM6go_W+yQ3|`}>R1Ifmq}xyc!?X{Je+Lu zo{x*mh4T~&Mc4zJOaLqeSIrKPtm#G4LdoeWNnod|7!|+`9j6=f_B77!Km|uWu3ECK zNpHjB*P%{3XSka?7U2m3Fp&hxs2VvRKwUQdfYh(ywhv)7-AL6%1s_0I3zPi6nAGG& z?_UNU1*1_E?8O_1$M#+KpWpCVel=72&3JqLkHIGf^BhOvM=Uk)ybj>c|EH_va3a8> z6VfPpY(!$xiTQwfdbRONFm!) zus_4X!w1_jii*ATMm`B()LIQyx&v$JDt8`_K5z|@{9BTDRDGqQWyyZ`fGQ46UEgtk z&w-Wo@eteoeB;pMiRN%_@Vdcz>xChH_3~vB(C681CA3DKVGSDwYD(Ah(=sb_MeShH zRrVGr!?TV485(*d5PNJt=br-c+26|rORcsv;CMI2ZRAp2MU zD%fO-k0~la;SMJM)a96GDR7g}qbEhjBme;z)3n5KygvH}Z>13Img!mJ)d zvD7rTw}p^pOt0J%?`DK9q{ zI?4-j^uRB=46J{lqt@3OCcqjTJ)$pYd&sDQ2*!rlw{IU8FE4~Es^+Hb28uDlQ3?RG zoC$uE;Y_%v)xrJWjVRz8U0*6(dot8GN*qTFyVH~q2Q7e>l@|kPithL;4+QVbX`5SE z08+jMPCkQ>%_}v5C3a}ls8ti#^bGvB?~+}?GH2stXTOfbHtK27*{}+LTq;7TMQ~AX zo>AW0x0k)65D+CKqyYhsq`$-nSRSCpQpVV7dcuXYw#Zjs?sOK4BRJt!ZP8-%k{ptf z`e0SMEi_JDLFoFfzi@~sUV!4<{PME}>b|~47Bm&IFYcBhBFvln4g6+5BV@;Hx_@a7 zEfjBgG>78syhVt8H8Y|nyw&^7d5(b%O7W-pUCrD7alC0L7B`u}6}ZelZgtB$eu~?j zOnHOb!vE|*%Fn~CCK42^Okt_uyobcO50dv(dJ`_@U@u(dYkl^DO~_Fr#hXPUlf_pK zC)l3~GWSYfL(b|Fsa1;Ug0h4HncLleTS+M1H8xU&bfsd?h1C|SK0-6yVMw1?qzNrG z&fl7LQh#?ZY-7&JR*L+teu_6B1vDgQw#8#-dmxjLbnb^a5YF`~z^TfVvPw^N;%=@`3T zRJ3G(XW$MwNBXBO#8U5?I^17;gO;zQ`+(YD=E|$Q`hxXXbKN4It0bbZ+PfDhvb&0H zSlD`|npUV3uip;8p32h;&kfi6dIo&ibbpRKB&1{yi*Lke)Xo1)F{xLY`7HbpOCZm# zOg#RBd8>|JYIS&Zzm;mvm2lSYZtG4B_^CN~3Ql+mC@8i7lQvo%{l-7>JF$#Uv2=1o zU0ua%&WbhFe(90#qa`c8OXw-ore0;#W!px{co#j8I>b53WW*Y^=|iodDY11}&u}7b z;UdSkMXj?Knx=y~>N@jT3k4huf|_p)7!>S=>DGE09F66)A9?AJN%BMQ|AQ)@DZ?L%y5w@%=xgVJOI zxv53$g4L5EgVA6=0}#WbK@@rZeJES4uN$#Of<~GP7ME9F=1!|+>ZvLcVs~g+^-nvl zG=1=79`>;KYG|maX5g%m5s(^nq!RU;a0OPVatn9`$egw5^Rx%23uL!&eN)dR61 zGRXG?`qE0TmHeRHOPHrmI(o>{LSx^c((&;89&Iygb@H{l>V`enCNva|9??$tQ*L?n zeO)e+>TsT1Q;_#8_^Rih#2)4CeMY|gaYSFEqxvV7)y2R|M&1p6n&1Y6`AntNlkR_; zwVc*Ayn8N{wUgU*aGGR-W~b)S44J(p=;|g*d;Op&M}kZ(&v6-}ZKF3Pk#SwgALchU z2KjUFgF2zFPIxGf#L%d}%Q)ozbRDgL+0c%19keE$#vWl58pckYgi#fD8m$W{gx>X= zY(#ao{vl3XPf#i)X(1LXUu1-ikGBwJ-WI>yo>J1C#u|VXmyFlEzl-~o_$b|VT{(9` z;AQ)b%$iHm)wg;|1q8n9?TynM$ScdZ{+)J%NJ`aRi)&uFrKvQ(UcwPdoo|2zskL;2`5U43g70bM|ZrcGU(rL*_!k z)@}+k`%-*=SSj2-FFjp!aXzD_aEOdz}ROY)RzRf)fhcB7RagzmRVykV;aN_5&07oY!D!a!dInODfs0S z8{0K7))08f>-h}YW3-s7pk-S$HkhJp>$XnrN+RFTl!zlGIe9*55U&VfF(5mzM=owOVq7zF3RNSm1!y~ zKR-f8R|0Z?h~Ja|vWjsJU;CxX}B{$B^0OF^(p$bVK5-vnxezo&>4UZFIN*Vw^~At&cMZ+r444 z{`q6#$ zUb?7Qed0knRsz0^V|1)Oje%Umb2f-4ix&qyeOk=wN{@io$=j6G5e}(fJ`i0-m=7It z8tw+M8u5ZbCb|`hOOmDa-TFBNYL7ZR?$CJpF3RdDzeKm~-$=>@h)t8Aumy5U_vv z=bs-<3OvP`)WBF)k(Yn^D?Kj7Of#yv=_4ZfZ;5&lOC*y6f&;Hn)5u6&Cp$Ldg~vE) zeH}MwbbsATGb_v};JYqdmfg|OfvuChefxHjeAmMJBoD;)MPAwPnHg=`?%lhsCR-x* z|MACnxAS)gmzEDY_M~%6Lv=H{fza#lv&deYKo;8HSB2;{|F_)RT|Kk^izG208HFS? z?z5g_O$sD2TcgF)syr4o!BK)#Y%;8UthnG?bNEB(&562n>5?S#eq}z-6`GPRNOi0( z4Cw{Y^SySPtl14i8p&^bKUFb?FgA@^D*!{c#3l>~_QH%Zsw|>E_L^ zf2^)*K!2{^g53?X$do~x zFM-m$I@8_MZM=hb5W)j$QX-_*coHbbYoIEkxULHhxO(_R6(-y2qNLw_e%%ajI zLZzy^z{mkn{tWyF`mg(3zX}uy6?b>{nh_EhP(5@?k}6rQwlCv%yw%7Krp*%q=M_S* z>HuP}i#2P4ew&_1(6>8PL1O9)oCWzwB2`V)LYBj!f(BI%(p`B9X-TL;H{dw1ZiP)5LG)s8?xMFLyp*w^2$Pt_RB`qUg4QFTP{nXUzh>P_#h)kN+LRJvB5}e%!dk+?tc`sB<72Dx2)qtoD21*vR*JaZw;LTdF@s#07os8{tf zo+upMuJ2djMH|8NrG}F#QLPqs{rT_y_~Q?L7VPwC(h`Ip!3cRjJlXEQRn!SCYrRDg zehtz$C*YJoBVd8sCjyLNzy{s`OCft|XoPj9Dln32;ARl3vXuP(b{CxRx&fZd=})jy z>b=I1Ap1c~Bb*qq9MLi~bP$9yL;a=pQS}vYY~>>255aSoLMy!i@%s8Du;0E#+mc_g znZgMgddaO8f%X$p-o#@6Sl}3Av$M8P@^Q6`aM(!%)VQaBO5lM`z{qD5xJ=v%oRwVq zKzHNn)%(UL0$zlNTi(0`LsN8MnEo=^q=m)M78col8XEBV!oacd+a`^G|Di)12gA?( zMA}G|M;;B6X#Df%hXQ9`g@?}q`#~0zVzlVkcU>bi8of&BnT^ifECR>oED#6KaPxYd zSw)>Lrzyqt=J}It1If@Ne&rfCIHsce4uDaF+hDP*fq<}Xxl;mU3ezF@qM@Y)(GT%x zI!4AUIo^s7W6~ggLLh7ZzI~k__Xt_EoPZQXub(2AR99tWVxCc5CZEFB+d_KkvfOJP zI1E;on3OazHMKrIoUJ1e*bXI+fG+) zLZ#5WIeMcOfzNIn$j(qx!7gb=K5F1Je@qeTiBl^;^+poMe7fsC^^o+|6>MIQHnBvOW5M zmH_D`z}HVeVLy80NaD!mqzKx32vP*oy?Ty4@ZNbySQlm#&YrwvGjRFDjw}V}^qH{r z;o#b>&{)r}`w;F5iM+qm^c5%uYV(6-;cga8)>KGSXsQ2&7?1F(dvz z9RBI~E7Az6JbJ1c!csOL9vw#P>Af^e>A*y~#yTsZq{`_ + + + +Weight of Evidence and Information Value +---------------------------------------- + +A common extension of the WoE is the information value (IV), which is a measure of the predictive power of a variable. The IV is calculated as follows: + +.. math:: + + IV = \sum_{i=1}^{n} (p_{i} - q_{i}) \cdot WoE_{i} + +Where, p_{i} is the percentage of positive cases in the i-th category, q_{i} is the percentage of negative cases in the i-th category, and WoE_{i} is the weight of evidence of the i-th category. + +The IV is a measure of the predictive power of a variable. The higher the IV value, the more predictive the variable is. So the combination of WoE with information value can be used for feature selection for binary classification problems. + + +Weight of Evidence and Information Value within Feature-engine +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +If you're asking yourself whether feature_engine allows you to automate this process, the answer is: of course!. You +can utilize the :class:`SelectByInformationValue()` class and it will handle all these steps for you. Again, +remember the given considerations. + +References +---------- + +- `Weight of Evidence: A Review of Concept and Methods `_ +- `Comparison and evaluation of landslide susceptibility maps obtained from weight of evidence, logistic regression, and artificial neural network models `_ +- `Can Weight of Evidence, Quantitative Bias, and Bounding Methods Evaluate Robustness of Real-World Evidence for Regulator and Health Technology Assessment Decisions on Medical Interventions`_ -In credit scoring, continuous variables are also transformed using the WoE. To do -this, first variables are sorted into a discrete number of bins, and then these -bins are encoded with the WoE as explained here for categorical variables. You can -do this by combining the use of the equal width, equal frequency or arbitrary -discretisers. Additional resources -------------------- From f4934dd712dafb02ae0484e7dab39923bd8e4665 Mon Sep 17 00:00:00 2001 From: hectorpatino Date: Mon, 15 Jul 2024 11:13:21 -0500 Subject: [PATCH 2/5] WoE User guide modifications-approved --- docs/user_guide/encoding/WoEEncoder.rst | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/docs/user_guide/encoding/WoEEncoder.rst b/docs/user_guide/encoding/WoEEncoder.rst index 632d45edf..24a7c90f2 100644 --- a/docs/user_guide/encoding/WoEEncoder.rst +++ b/docs/user_guide/encoding/WoEEncoder.rst @@ -335,6 +335,7 @@ Finally, we can visualize the values of the WoE encoded variables respect to the plt.grid(axis='y') plt.show() +In the following plot, we can see the WoE for different categories of the variable 'age': .. figure:: ../../images/woe_encoding.png :width: 600 @@ -343,7 +344,7 @@ Finally, we can visualize the values of the WoE encoded variables respect to the WoE for Age -We can visualize WoE for different categories of the variable 'age'. The WoE values are in the y-axis, and the categories are in the x-axis. We see that the WoE values are monotonically increasing, which is the expected behavior of the WoE. If we check the category 4 (which is a label), we can see the WoE is around -0.45 which means that it has a small portion of positive cases compared to negative cases. This means that this category has a low probability of survival. +The WoE values are in the y-axis, and the categories are in the x-axis. We see that the WoE values are monotonically increasing, which is the expected behavior of the WoE. If we check the category 4 (which is a label), we can see the WoE is around -0.45 which means that it has a small portion of positive cases compared to negative cases. This means that this category has a low probability of survival. Adding a model to the pipeline @@ -377,7 +378,7 @@ The accuracy of the model presented below: Accuracy: 0.76 -The accuracy of the model is 0.76, which is a good result for a first model. We can improve the model by tuning the hyperparameters of the logistic regression model or by using other models. Please note that accuracy may not be the best metric for this problem, as the dataset is imbalanced. We recommend using other metrics such as the F1 score, precision, recall, or the ROC-AUC score. You can learn more about imbalance datasets in our `course `_ +The accuracy of the model is 0.76, which is a good result for a first model. We can improve the model by tuning the hyperparameters of the logistic regression model. Please note that accuracy may not be the best metric for this problem, as the dataset is imbalanced. We recommend using other metrics such as the F1 score, precision, recall, or the ROC-AUC score. You can learn more about imbalance datasets in our `course `_ From 62264eddc04cc5c184c41554803649a28359b0a8 Mon Sep 17 00:00:00 2001 From: hectorpatino Date: Mon, 15 Jul 2024 11:30:44 -0500 Subject: [PATCH 3/5] WoE User guide modifications-approved-tox --- docs/user_guide/encoding/WoEEncoder.rst | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/user_guide/encoding/WoEEncoder.rst b/docs/user_guide/encoding/WoEEncoder.rst index 24a7c90f2..db835f5e5 100644 --- a/docs/user_guide/encoding/WoEEncoder.rst +++ b/docs/user_guide/encoding/WoEEncoder.rst @@ -408,7 +408,7 @@ References - `Weight of Evidence: A Review of Concept and Methods `_ - `Comparison and evaluation of landslide susceptibility maps obtained from weight of evidence, logistic regression, and artificial neural network models `_ -- `Can Weight of Evidence, Quantitative Bias, and Bounding Methods Evaluate Robustness of Real-World Evidence for Regulator and Health Technology Assessment Decisions on Medical Interventions`_ +- `Can weight of evidence, quantitative bias, and bounding methods evaluate robustness of real-world evidence for regulator and health technology assessment decisions on medical interventions `_ Additional resources From 44a9c38e6363a06c0521113cd75696994eff5e45 Mon Sep 17 00:00:00 2001 From: hectorpatino Date: Mon, 15 Jul 2024 11:59:26 -0500 Subject: [PATCH 4/5] refactor (discretizers): Due to compatibility issues with py312 and numpy 2.0 np.Inf is refactored to np.inf --- .../test_discretisation/test_arbitrary_discretiser.py | 10 +++++----- tests/test_discretisation/test_base_discretizer.py | 6 +++--- .../test_check_estimator_discretisers.py | 6 +++--- 3 files changed, 11 insertions(+), 11 deletions(-) diff --git a/tests/test_discretisation/test_arbitrary_discretiser.py b/tests/test_discretisation/test_arbitrary_discretiser.py index 2f8c37f82..44a020c04 100644 --- a/tests/test_discretisation/test_arbitrary_discretiser.py +++ b/tests/test_discretisation/test_arbitrary_discretiser.py @@ -13,19 +13,19 @@ def test_arbitrary_discretiser(): data = pd.DataFrame( california_dataset.data, columns=california_dataset.feature_names ) - user_dict = {"HouseAge": [0, 20, 40, 60, np.Inf]} + user_dict = {"HouseAge": [0, 20, 40, 60, np.inf]} data_t1 = data.copy() data_t2 = data.copy() # HouseAge is the median house age in the block group. data_t1["HouseAge"] = pd.cut( - data["HouseAge"], bins=[0, 20, 40, 60, np.Inf], include_lowest=True + data["HouseAge"], bins=[0, 20, 40, 60, np.inf], include_lowest=True ) data_t1["HouseAge"] = data_t1["HouseAge"].astype(str) data_t2["HouseAge"] = pd.cut( data["HouseAge"], - bins=[0, 20, 40, 60, np.Inf], + bins=[0, 20, 40, 60, np.inf], labels=False, include_lowest=True, ) @@ -53,7 +53,7 @@ def test_arbitrary_discretiser(): def test_error_if_input_df_contains_na_in_transform(df_vartypes, df_na): # test case 1: when dataset contains na, transform method - age_dict = {"Age": [0, 10, 20, 30, np.Inf]} + age_dict = {"Age": [0, 10, 20, 30, np.inf]} with pytest.raises(ValueError): transformer = ArbitraryDiscretiser(binning_dict=age_dict) @@ -119,6 +119,6 @@ def test_error_when_nan_introduced_during_transform(): def test_error_if_not_permitted_value_is_errors(): - age_dict = {"Age": [0, 10, 20, 30, np.Inf]} + age_dict = {"Age": [0, 10, 20, 30, np.inf]} with pytest.raises(ValueError): ArbitraryDiscretiser(binning_dict=age_dict, errors="medialuna") diff --git a/tests/test_discretisation/test_base_discretizer.py b/tests/test_discretisation/test_base_discretizer.py index 1a27b8f12..fc8110ff1 100644 --- a/tests/test_discretisation/test_base_discretizer.py +++ b/tests/test_discretisation/test_base_discretizer.py @@ -43,7 +43,7 @@ def fit(self, X): california_dataset.data, columns=california_dataset.feature_names ) self.variables_ = ["HouseAge"] - self.binner_dict_ = {"HouseAge": [0, 20, 40, 60, np.Inf]} + self.binner_dict_ = {"HouseAge": [0, 20, 40, 60, np.inf]} self.n_features_in_ = data.shape[1] self.feature_names_in_ = california_dataset.feature_names return self @@ -60,12 +60,12 @@ def test_transform(): # HouseAge is the median house age in the block group. data_t1["HouseAge"] = pd.cut( - data["HouseAge"], bins=[0, 20, 40, 60, np.Inf], include_lowest=True + data["HouseAge"], bins=[0, 20, 40, 60, np.inf], include_lowest=True ) data_t1["HouseAge"] = data_t1["HouseAge"].astype(str) data_t2["HouseAge"] = pd.cut( data["HouseAge"], - bins=[0, 20, 40, 60, np.Inf], + bins=[0, 20, 40, 60, np.inf], labels=False, include_lowest=True, ) diff --git a/tests/test_discretisation/test_check_estimator_discretisers.py b/tests/test_discretisation/test_check_estimator_discretisers.py index 749d74561..c6a4951ff 100644 --- a/tests/test_discretisation/test_check_estimator_discretisers.py +++ b/tests/test_discretisation/test_check_estimator_discretisers.py @@ -17,7 +17,7 @@ DecisionTreeDiscretiser(regression=False), EqualFrequencyDiscretiser(), EqualWidthDiscretiser(), - ArbitraryDiscretiser(binning_dict={"x0": [-np.Inf, 0, np.Inf]}), + ArbitraryDiscretiser(binning_dict={"x0": [-np.inf, 0, np.inf]}), GeometricWidthDiscretiser(), ] @@ -30,14 +30,14 @@ def test_check_estimator_from_sklearn(estimator): @pytest.mark.parametrize("estimator", _estimators) def test_check_estimator_from_feature_engine(estimator): if estimator.__class__.__name__ == "ArbitraryDiscretiser": - estimator.set_params(binning_dict={"var_1": [-np.Inf, 0, np.Inf]}) + estimator.set_params(binning_dict={"var_1": [-np.inf, 0, np.inf]}) return check_feature_engine_estimator(estimator) @pytest.mark.parametrize("transformer", _estimators) def test_transformers_within_pipeline(transformer): if transformer.__class__.__name__ == "ArbitraryDiscretiser": - transformer.set_params(binning_dict={"feature_1": [-np.Inf, 0, np.Inf]}) + transformer.set_params(binning_dict={"feature_1": [-np.inf, 0, np.inf]}) X = pd.DataFrame({"feature_1": [1, 2, 3, 4, 5], "feature_2": [6, 7, 8, 9, 10]}) y = pd.Series([0, 1, 0, 1, 0]) From 5804f3609e6059606d9b05d58b1eab6a6fffbbf9 Mon Sep 17 00:00:00 2001 From: Soledad Galli Date: Tue, 16 Jul 2024 09:09:07 +0200 Subject: [PATCH 5/5] edit copy --- docs/user_guide/encoding/WoEEncoder.rst | 183 +++++++++++++++++------- 1 file changed, 133 insertions(+), 50 deletions(-) diff --git a/docs/user_guide/encoding/WoEEncoder.rst b/docs/user_guide/encoding/WoEEncoder.rst index db835f5e5..3494d5a21 100644 --- a/docs/user_guide/encoding/WoEEncoder.rst +++ b/docs/user_guide/encoding/WoEEncoder.rst @@ -5,11 +5,17 @@ Weight of Evidence (WoE) ======================== -The term Weight of Evidence (WoE) can be traced to the financial sector, especially to 1983, when it took on an important role in describing the key components of credit risk analysis and credit scoring. Since then, it has been used for medical research, GIS studies, and more (see references below for review). +The term Weight of Evidence (WoE) can be traced to the financial sector, especially to +1983, when it took on an important role in describing the key components of credit risk +analysis and credit scoring. Since then, it has been used for medical research, GIS +studies, and more (see references below for review). -The WoE is a statistical data-driven method based on Bayes' theorem and the concepts of prior and posterior probability, so the concepts of log odds, events, and non-events are crucial to understanding how the weight of evidence works. +The WoE is a statistical data-driven method based on Bayes' theorem and the concepts of +prior and posterior probability, so the concepts of log odds, events, and non-events +are crucial to understanding how the weight of evidence works. -The WoE is only defined for binary classification problems. In other words, we can only encode variables using the WoE when the target variable is binary. +The WoE is only defined for binary classification problems. In other words, we can only +encode variables using the WoE when the target variable is binary. Formulation ----------- @@ -25,45 +31,54 @@ We discuss the formula in the next section. Calculation ----------- -We have a dataset with a binary dependent variable with two categories, 0 and 1, and -a categorical predictor variable named variable A with three categories (A1, A2, and A3). The dataset has the following characteristics: +How is the WoE calculated? Let's say we have a dataset with a binary dependent variable +with two categories, 0 and 1, and a categorical predictor variable named variable A +with three categories (A1, A2, and A3). The dataset has the following characteristics: - There are 20 positive (1) cases and 80 negative (0) cases in the target variable. - Category A1 has 10 positive cases and 15 negative cases. - Category A2 has 5 positive cases and 15 negative cases. - Category A3 has 5 positive cases and 50 negative cases. -First, we group the data by each category of the predictor variable where the target is positive, and then we divide by the total number of positive cases. Then we do the same for the negative cases: +First, we find out the number of instances with a positive target value (1) per category, +and then we divide that by the total number of positive cases in the data. Then we determine +the number of instances with target value of 0 per category and divide that by the total +number of negative instances in the dataset: - For category A1, we have 10 positive cases and 15 negative cases, resulting in a positive ratio of 10/20 and a negative ratio of 15/80. This means that the positive ratio is 0.5 and the negative ratio is 0.1875. - For category A2, we have 5 positive cases out of 20 positive cases, giving us a positive ratio of 5/20 and a negative ratio of 15/80. This results in a positive ratio of 0.25 and a negative ratio of 0.1875. - For category A3, we have 5 positive cases out of 20 positive cases, resulting in a positive ratio of 5/20, and a 50/80 negative ratio. So the positive ratio is 0.25, and the negative ratio is 0.625. -Now we calculate the log of the ratio of the percentage of positive cases in each category. +Now we calculate the log of the ratio of positive cases in each category: - For category A1, we have log (0.5/ 0.1875) = 0.98. - For category A2, we have log (0.25/ 0.1875) = 0.28. - For category A3, we have log (0.25/0.625) =-0.91. -Finally, we replace the categories (A1, A2, and A3) of the independent variable Variable A with the WoE values (0.98, 0.28, -0.91). +Finally, we replace the categories (A1, A2, and A3) of the independent variable A with +the WoE values: 0.98, 0.28, -0.91. Characteristics of the WoE -------------------------- -The beauty of the WoE, is that we can directly understand the impact of the category on the probability of success (target variable being 1): +The beauty of the WoE, is that we can directly understand the impact of the category on +the probability of success (target variable being 1): - If WoE values are negative, there are more negative cases than positive cases for the category. - If WoE values are positive, there are more positive cases than negative cases for that category. - If WoE is 0, then there is an equal number of positive and negative cases for that category. -In other words, for categories with positive WoE, the probability of success is high, for categories with negative WoE, the probability of success is low, and for those with WoE of zero, there are equal chances for both target outcomes. +In other words, for categories with positive WoE, the probability of success is high, +for categories with negative WoE, the probability of success is low, and for those with +WoE of zero, there are equal chances for both target outcomes. Advantages of the WoE --------------------- -In addition to the intuitive interpretation of the WoE valuess, the WoE shows the following advantages: +In addition to the intuitive interpretation of the WoE values, the WoE shows the following +advantages: - It creates monotonic relationships between the encoded variable and the target. - It returns numeric variables on a similar scale. @@ -72,49 +87,76 @@ In addition to the intuitive interpretation of the WoE valuess, the WoE shows th Uses of the WoE --------------- -In general, we use the WoE to encode both categorical and numerical variables. For continuous variables, we first need to do binning, that is, sort the variables into discrete intervals. You can do this by preprocessing the variable using any of Feature-engine's discretizers. +In general, we use the WoE to encode both categorical and numerical variables. For +continuous variables, we first need to do binning, that is, sort the variables into +discrete intervals. You can do this by preprocessing the variable using any of +Feature-engine's discretizers. -Some authors have extended the Weight of Evidence approach to neural networks and other algorithms, and although they have shown good results, the predictive modeling performance of Weight of Evidence was superior when used with -logistic regression models (see reference below). +Some authors have extended the Weight of Evidence approach to neural networks and other +algorithms, and although they have shown good results, the predictive modeling performance +of Weight of Evidence was superior when used with logistic regression models (see +reference below). Limitations of the WoE ---------------------- -As the methodology to calculate the WoE is based on ratios and logarithm, the value is not defined when `p(X=xj|Y = 1) = 0` or `p(X=xj|Y=0) = 0`. For the latter, the division by 0 is not defined, and for the former, the log of 0 is not defined. +As the methodology to calculate the WoE is based on ratios and logarithm, the WoE value +is not defined when `p(X=xj|Y = 1) = 0` or `p(X=xj|Y=0) = 0`. For the latter, the division +by 0 is not defined, and for the former, the log of 0 is not defined. -This occurs when a category shows only 1 of the possible values of the target (either it always takes 1 or 0). In practice, this happens mostly when a category has a low frequency in the dataset, that is, when only very few observations show that category. +This occurs when a category shows only 1 of the possible values of the target (either it +always takes 1 or 0). In practice, this happens mostly when a category has a low frequency +in the dataset, that is, when only very few observations show that category. -To overcome this limitation, consider using a variable transformation method to group those categories together, for example by using Feature-engine's :class:`RareLabelEncoder()`. +To overcome this limitation, consider using a variable transformation method to group +those categories together, for example by using Feature-engine's :class:`RareLabelEncoder()`. - -Taking into account the above considerations, conducting a detailed exploratory data analysis (EDA) is essential as -part of the data science and model-building process. Integrating these considerations and practices not only -enhances the feature engineering process but also improves the performance of your models. +Taking into account the above considerations, conducting a detailed exploratory data +analysis (EDA) is essential as part of the data science and model-building process. +Integrating these considerations and practices not only enhances the feature engineering +process but also improves the performance of your models. Unseen categories ----------------- -When using the WoE, we define the mappings, that is, the WoE values per category using the observations from the training set. If the test set shows new (unseen) categories, we'll lack a WoE value for them, and won't be able to encode them. +When using the WoE, we define the mappings, that is, the WoE values per category using +the observations from the training set. If the test set shows new (unseen) categories, +we'll lack a WoE value for them, and won't be able to encode them. -This is a known issue, without an elegant solution. If the new values appear in continuous variables, consider changing the size and number of the intervals. If the unseen categories are seen in categorical variables, consider grouping low frequency categories before doing the encoding. +This is a known issue, without an elegant solution. If the new values appear in continuous +variables, consider changing the size and number of the intervals. If the unseen categories +appear in categorical variables, consider grouping low frequency categories before doing +the encoding. WoEEncoder ---------- -The :class:`WoEEncoder()` allows you to automate the process of calculating weight of evidence for a given set of features. By default, :class:`WoEEncoder()` will encode all categorical variables. You can encode just a subset by passing the variables names in a list to the `variables` parameter. +The :class:`WoEEncoder()` allows you to automate the process of calculating weight of +evidence for a given set of features. By default, :class:`WoEEncoder()` will encode all +categorical variables. You can encode just a subset by passing the variables names in a +list to the `variables` parameter. -By default, :class:`WoEEncoder()` will not encode numerical variables, instead, it will raise an error. If you want to encode numerical, for example discrete variables, set `ignore_format` to `True`. +By default, :class:`WoEEncoder()` will not encode numerical variables, instead, it will +raise an error. If you want to encode numerical, for example discrete variables, set +`ignore_format` to `True`. -:class:`WoEEncoder()` does not handle missing values automatically, so make sure to replace them with a suitable value before the encoding. You can impute missing values with Feature-engine's imputers. +:class:`WoEEncoder()` does not handle missing values automatically, so make sure to +replace them with a suitable value before the encoding. You can impute missing values +with Feature-engine's imputers. -:class:`WoEEncoder()` will ignore unseen categories by default, in which case, they will be replaced by np.nan after the encoding. You have the option to make the encoder raise an error instead, by setting `unseen='raise'`. You can also replace unseen categories by an arbitrary value you need to define in `fill_value`, although we do not recommend that option. +:class:`WoEEncoder()` will ignore unseen categories by default, in which case, they will +be replaced by np.nan after the encoding. You have the option to make the encoder raise +an error instead, by setting `unseen='raise'`. You can also replace unseen categories +by an arbitrary value you need to define in `fill_value`, although we do not recommend +this option because it may lead to unpredictable results. Python example -------------- -In the rest of the document, we'll show :class:`WoEEncoder()`'s functionality. Let's look at an example using the Titanic Dataset. +In the rest of the document, we'll show :class:`WoEEncoder()`'s functionality. Let's +look at an example using the Titanic Dataset. First, let's load the data and separate the dataset into train and test: @@ -149,7 +191,8 @@ We see the resulting dataframe below: 686 3 female 22.000000 0 0 7.7250 M Q Before we encode the variables, we group infrequent categories into one -category, which we'll call 'Rare'. For this, we use the :class:`RareLabelEncoder()` as follows: +category, which we'll call 'Rare'. For this, we use the :class:`RareLabelEncoder()` as +follows: .. code:: python @@ -181,14 +224,15 @@ evidence, only in the 3 indicated variables: # fit the encoder woe_encoder.fit(train_t, y_train) -With `fit()` the encoder learns the weight of the evidence for each category, which are stored in its `encoder_dict_` parameter: +With `fit()` the encoder learns the weight of the evidence for each category, which are +stored in its `encoder_dict_` parameter: .. code:: python woe_encoder.encoder_dict_ In the `encoder_dict_` we find the WoE for each one of the categories of the -variables to encode. This way, we can map the original values to the new value: +variables to encode. This way, we can map the original values to the new values: .. code:: python @@ -209,7 +253,8 @@ Now, we can go ahead and encode the variables: print(train_t.head()) -Below we see the resulting dataset with the weight of the evidence replacing the original variable values: +Below we see the resulting dataset with the weight of the evidence replacing the original +variable values: .. code:: python @@ -226,11 +271,12 @@ WoE in categorical and numerical variables In the previous example, we encoded only the variables 'cabin', 'pclass', 'embarked', and left the rest of the variables untouched. In the following example, we will use -Feature-engine's pipeline to transform variables in sequence. We'll group rare categories in categorical variables. Next, we'll discretize numerical variables. And finally, we'll encode them all with the WoE. +Feature-engine's pipeline to transform variables in sequence. We'll group rare categories +in categorical variables. Next, we'll discretize numerical variables. And finally, we'll +encode them all with the WoE. First, let's load the data and separate it into train and test: - .. code:: python from sklearn.model_selection import train_test_split @@ -272,7 +318,9 @@ Let's define lists with the categorical and numerical variables: numerical_features = ['fare', 'age'] all = categorical_features + numerical_features -Now, we will set up the pipeline to first discretize the numerical variables, then group rare labels and low frequency intervals into a common group, and finally encode all variables with the WoE: +Now, we will set up the pipeline to first discretize the numerical variables, then group +rare labels and low frequency intervals into a common group, and finally encode all +variables with the WoE: .. code::: python @@ -289,7 +337,8 @@ We have created a variable transformation pipeline with the following steps: - Next, we use :class:`RareLabelEncoder()` to group infrequent categories and intervals into one group. - Finally, we use the :class:`WoEEncoder()` to replace values in all variables with the weight of the evidence. -Now, we can go ahead and fit the pipeline to the train set so that the different transformers learn the parameters for the variable transformation. +Now, we can go ahead and fit the pipeline to the train set so that the different +transformers learn the parameters for the variable transformation. .. code:: python @@ -316,7 +365,9 @@ We see the resulting dataframe below: 1193 0.012075 686 0.012075 -Finally, we can visualize the values of the WoE encoded variables respect to the original values to corroborate the sigmoid function shape, which is the expected behavior of the WoE: +Finally, we can visualize the values of the WoE encoded variables respect to the original +values to corroborate the sigmoid function shape, which is the expected behavior of the +WoE: .. code:: python @@ -335,22 +386,45 @@ Finally, we can visualize the values of the WoE encoded variables respect to the plt.grid(axis='y') plt.show() -In the following plot, we can see the WoE for different categories of the variable 'age': +In the following plot, we can see the WoE for different categories of the variable +'age': .. figure:: ../../images/woe_encoding.png :width: 600 :figclass: align-center :align: left - WoE for Age +| +| +| +| +| +| +| +| +| +| +| +| +| +| +| +| +| -The WoE values are in the y-axis, and the categories are in the x-axis. We see that the WoE values are monotonically increasing, which is the expected behavior of the WoE. If we check the category 4 (which is a label), we can see the WoE is around -0.45 which means that it has a small portion of positive cases compared to negative cases. This means that this category has a low probability of survival. +The WoE values are in the y-axis, and the categories are in the x-axis. We see that the +WoE values are monotonically increasing, which is the expected behavior of the WoE. If +we look at category 4, we can see the WoE is around -0.45 which means that in this age +bracket there was a small portion of positive cases (people who survived) compared to +negative cases (non-survivors). In other words, people within this age interval had +a low probability of survival. Adding a model to the pipeline ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -To complete the demo, we can add a logistic regression model to the pipeline to obtain predictions of survival after the variable transformation. +To complete the demo, we can add a logistic regression model to the pipeline to obtain +predictions of survival after the variable transformation. .. code:: python @@ -366,42 +440,51 @@ To complete the demo, we can add a logistic regression model to the pipeline to ]) - pipe.fit(X_train, y_train) pipe.fit(X_train, y_train) y_pred = pipe.predict(X_test) accuracy = accuracy_score(y_test, y_pred) print(f"Accuracy: {accuracy:.2f}") -The accuracy of the model presented below: +The accuracy of the model is shown below: .. code:: python Accuracy: 0.76 -The accuracy of the model is 0.76, which is a good result for a first model. We can improve the model by tuning the hyperparameters of the logistic regression model. Please note that accuracy may not be the best metric for this problem, as the dataset is imbalanced. We recommend using other metrics such as the F1 score, precision, recall, or the ROC-AUC score. You can learn more about imbalance datasets in our `course `_ +The accuracy of the model is 0.76, which is a good result for a first model. We can +improve the model by tuning the hyperparameters of the logistic regression model. Please +note that accuracy may not be the best metric for this problem, as the dataset is +imbalanced. We recommend using other metrics such as the F1 score, precision, recall, or +the ROC-AUC score. You can learn more about imbalance datasets in our +`course `_. Weight of Evidence and Information Value ---------------------------------------- -A common extension of the WoE is the information value (IV), which is a measure of the predictive power of a variable. The IV is calculated as follows: +A common extension of the WoE is the information value (IV), which is a measure of the +predictive power of a variable. The IV is calculated as follows: .. math:: IV = \sum_{i=1}^{n} (p_{i} - q_{i}) \cdot WoE_{i} -Where, p_{i} is the percentage of positive cases in the i-th category, q_{i} is the percentage of negative cases in the i-th category, and WoE_{i} is the weight of evidence of the i-th category. +Where, `pi` is the percentage of positive cases in the i-th category, `qi` is the +percentage of negative cases in the i-th category, and WoE_{i} is the weight of evidence +of the i-th category. -The IV is a measure of the predictive power of a variable. The higher the IV value, the more predictive the variable is. So the combination of WoE with information value can be used for feature selection for binary classification problems. +The IV is a measure of the predictive power of a variable. The higher the IV value, the +more predictive the variable is. So the combination of WoE with information value can be +used for feature selection for binary classification problems. Weight of Evidence and Information Value within Feature-engine ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ -If you're asking yourself whether feature_engine allows you to automate this process, the answer is: of course!. You -can utilize the :class:`SelectByInformationValue()` class and it will handle all these steps for you. Again, -remember the given considerations. +If you're asking yourself whether Feature-engine allows you to automate this process, +the answer is: of course! You can utilize the :class:`SelectByInformationValue()` class +and it will handle all these steps for you. Again, remember the given considerations. References ----------