;i  ddlZn"#e$rGddZeZYnwxYw ddlmZddlZejejr edddl m Z n#e$rddlZej sddlm Z YnwxYwddl Z ddl Z ddlZddlZddlmZdd lmZdd lmZd d gZe je jejd ZejdRdededefdZej sddlmZdZ e fdZ!dZ"dZ#dZ$dZ%ejde&fdZ'dZ(dZ)dZ*dSd Z+d!e&fd"Z,ejd#edefd$Z-ej sd#edefd%Z-Gd&d'Z.Gd(d)Z/d!e&fd*Z0d+e&fd,Z1ejde&fd-Z2Gd.d/eZ3Gd0d1e3Z4Gd2d3e3Z5dTd5Z6dTd6Z7ej8d7ej9ej:zj;Z<ej8d8ej9ej:zj;Z=ej8d9ej9ej:zj>Z?d:Z@d;ZAd<ZBejCeDejEZEejCeDeDgd=ZFejCeDeDgd>ZGejCeeeHeFeGzZIdSd?ZJej8d@ejKjLZMdAZNej8dBjOZPdCZQdDZRdEZSdFZTdGZUdHZVdSdIZWejdJZXejdKZYejdLZZdMZ[ej\ej]GdNdOe Z^e_dPkrddQlm`Z`e`jadSdS)UNc*eZdZdZdZdZdZdZdS) fake_cythonFc|SNselffuncs BC:\PYTHON\MyICR_Workspace\venv\Lib\site-packages\lxml/html/diff.pycfunczfake_cython.cfuncsd{c|Srrrs r cclasszfake_cython.cclass st r c|Srr)r _values r declarezfake_cython.declare sE\r cdS)Nobjectr)r type_names r __getattr__zfake_cython.__getattr__ sr N)__name__ __module__ __qualname__compiledr rrrrr r rrs9***+++11199999r r)_difflibzLEmbedded difflib is not compiled to a fast binary, using the stdlib instead.)SequenceMatcher)etree)fragment_fromstring)defs html_annotatehtmldiff)keyz&z<z>z"z'text_escapesreturnc\dgdz}|D]f}|dxx|dkzcc<|dxx|dkzcc<|dxx|dkzcc<|d xx|d kzcc<|d xx|d kzcc<gtdD],}||r"|d |||}-|S)NFr&r<>"'z&<>"')rangereplace)r&r'r4chis r html_escaper7-s %gkG   bCi  bCi  bCi  bCi  bCi 1XX:: 1: :<< Xa[99D Kr )escapec.dt|d|dS)Nz zr7)r&versions r default_markupr<Es%Gddd $$r cd|D}|d}|ddD]}t|||}t|}t||}d|S)a doclist should be ordered from oldest to newest, like:: >>> version1 = 'Hello World' >>> version2 = 'Goodbye World' >>> print(html_annotate([(version1, 'version 1'), ... (version2, 'version 2')])) Goodbye World The documents must be *fragments* (str/UTF8 or unicode), not complete documents The markup argument is a function to markup the spans of words. This function is called like markup('Hello', 'version 2'), and returns HTML. The first argument is text and never includes any markup. The default uses a span with a title: >>> print(default_markup('Some Text', 'by Joe')) Some Text c4g|]\}}t||Sr)tokenize_annotated).0docr;s r z!html_annotate..cs6...!S'$C11...r rrN)html_annotate_merge_annotationscompress_tokensmarkup_serialize_tokensjoinstrip)doclistmarkup tokenlist cur_tokenstokensresults r r"r"Is4..%,...I1JABB-' F;;; !,,J $Z 8 8F 776?? " ""r c@t|d}|D] }||_ |S)zFTokenize a document and add an annotation attribute to each token F include_hrefs)tokenize annotation)rArSrMtoks r r?r?qs3c / / /F$$# Mr ct||}|}|D]2\}}}}}|dkr$|||} |||} t| | 3dS)zMerge the annotations from tokens_old into tokens_new, when the tokens in the new document already existed in the old document. abequalN)InsensitiveSequenceMatcher get_opcodescopy_annotations) tokens_old tokens_newscommandscommandi1i2j1j2eq_oldeq_news r rDrDys #Z:>>>A}}H#+--RR g  2&F2&F VV , , , --r ct|t|ksJt||D]\}}|j|_dS)zN Copy annotations from the tokens listed in src to the tokens in dest N)lenziprS)srcdestsrc_tokdest_toks r r\r\sU s88s4yy  d^^11%011r c|dg}|ddD]R}|js4|djs'|dj|jkrt||=||S|S)zl Combine adjacent tokens when there is no HTML between the tokens, and they share an annotation rrN)pre_tags post_tagsrScompress_merge_backappend)rMrNrTs r rErEs Qi[Fabbz  2J( r %77  , , , , MM#     Mr rMc|d}t|tust|tur||dS||jz|z}t||j|j|j}|j|_||d<dS)zY Merge tok into the last element of tokens (modifying the list of tokens in-place). rprqrrtrailing_whitespaceN)typetokenrtrwrqrrrS)rMrTlastr&mergeds r rsrss ":D Dzz$s))5"8"8 cd..4t $ !$+.+BDDD!Or r c#K|D]M}|jEd{V|}|||j|jz}|V|jEd{VNdS)zz Serialize the list of tokens into a list of text chunks, calling markup_func around text to add annotations. N)rqhtmlrSrwrr)rM markup_funcryr}s r rFrFs ##>!!!!!!!zz||{4!122U5NN ?"""""""" ##r c,t|}t|}t||} d|}n/#tt f$r}t |d}Yd}~nd}~wwxYwt|S)a Do a diff of the old and new document. The documents are HTML *fragments* (str/UTF8 or unicode), they are not complete documents (i.e., no tag). Returns HTML with and tags added around the appropriate text. Markup is generally ignored, with the markup from new_html preserved, and possibly some markup from old_html (though it is considered acceptable to lose some of the old markup). Only the words in the HTML are diffed. The exception is tags, which are treated like words, and the href attribute of tags, which are noted inside the tag itself when there are changes. rCN)rRhtmldiff_tokensrGrH ValueError TypeErrorprintfixup_ins_del_tags)old_htmlnew_htmlold_html_tokensnew_html_tokensrNexcs r r#r#s"x((Ox((O _o > >F&&((  " c  f % %%s'AB)A??Bct||}|}g}|D]\}}}}} |dkr-|t||| d;|dks|dkr't||| } t | ||dks|dkr't|||} t | |t ||S)z] Does a diff on the tokens themselves, returning a list of text chunks (not tokens). rVrYT)rYinsertr4delete)rZr[extend expand_tokens merge_insert merge_deletecleanup_delete) html1_tokens html2_tokensr_r`rNrarbrcrdre ins_tokens del_tokenss r rrs" #\\BBBA}}H F#+ - -RR g   MM- RU(;4HHH I I I  h  'Y"6"6&|BrE':;;J V , , , h  'Y"6"6&|BrE':;;J V , , , 6 Mr Fc#K|D]C}|jEd{V|r|js||jzV|jEd{VDdS)zeGiven a list of tokens, return a generator of the chunks of text for the data in the tokens. N)rqhide_when_equalr}rwrr)rMrYrys r rrs##>!!!!!!! ;E1 ;**,,!:: : : :?"""""""" ##r rActt|D]\}}d|D}|dkr|r+|dds|dxxdz cc<|d|||ddr|ddd|d<|d||dS)z| doc is the already-handled document (as a list of text chunks); here we add ins_chunks to the end of that. cg|] }|d S)rr)r@items r rBz merge_insert..s444d$q'444r rXrp zNz )group_by_first_itemmark_unbalancedendswithrtr) ins_chunksrAbalanced marked_chunkschunkss r rr s$7z7R7R#S#S-44m444 s?? 3r7++C00 B3 JJw    JJv   2w$$ 'b'#2#,B JJy ! ! ! ! JJv    r chunkc|ddkrdSd}t|D]@\}}|dkrd}|dkr |||cS|r |||cSA||dS)Nrr,rCr/r-r.) enumerateisspace)r start_posr6r5s r tag_name_of_chunkr+s  Qx3rI5!!&&2 99II 3YY1% % % % ZZ\\ &1% % % % &  r c`|ddddS)Nrrz<>/)splitrH)rs r rr@s){{4##A&,,U333r ceZdZdS) DEL_STARTNrrrrr r rrGDr rceZdZdS)DEL_ENDNrrr r rrIrr rc|t|||tdS)z Adds the text chunks in del_chunks to the document doc (another list of text chunks) with marker to show it is a delete. cleanup_delete later resolves these markers into tags.N)rtrrr) del_chunksrAs r rrMs@ JJyJJzJJwr rct|}d} |t|}|t|dz}n#t$rYd SwxYwdx}}dx}}t ||dz|} | D]\} } | dkrn|dz }t | } t|dz|D]l} || turnZ|| }|ddks |ddkrn8t |}|dkrn!|dks Jd||| krn|dz }mt| D]\} } | d krn|dz }t | }t|dz d d D]_} || turnM|| }|ddkr|ddkrn+t |}|dks|dkrn||krn|dz }` ||z }t|dz||zdzD]} || ||<|dz }|r1||dz  d s||dz xxd z cc<d ||<|dz }t||zdz||z D]} || ||<|dz }||dz  d r||dz d d ||dz <d||<|dz }t||z |D]} || ||<|dz }||||zdz=|})a Cleans up any DEL_START/DEL_END markers in the document, replacing them with . To do this while keeping the document valid, it may need to drop some tags (either start or end tags). It may also move the del into adjacent tags to try to move it to a similar location where it was originally located (e.g., moving a delete into preceding
tag, if the del looks like (DEL_START, 'Text
', DEL_END) rrusr,rinsdelzUnexpected delete tag: uerprzNz ) riindexrrrrrr3reversedr)r chunk_countr del_startdel_endshift_end_leftshift_start_rightunbalanced_startunbalanced_enddeleted_chunksr del_chunkunbalanced_start_namer6rnameunbalanced_end_nameposs r rrWsf++KIp ; Y ::I ll7IM::GG     EE  ./.*,-->( ! G0C)DEE$2 ' ' Hi4  ! $5i$@$@ !719k22 ' '!9 ))Eq 8s??eAh#ooE(//5==Eu}}}&I&I&I}}}000E!Q&!!$,N#;#; $ $ Hi4 a N"3I">"> 9q="b11 $ $!9''Eq 8s??uQx3E(//5==DEMME...E!# ,.(y1}i2C&Ca&GHH  A )F3K 1HCC  #vcAg//44 # 37OOOs "OOOs  qy#33a7>9QRR  A )F3K 1HCC #'? # #C ( ( 3$S1Wocrc2F37Os  qw/99  A )F3K 1HCC 3#44q88 9 apsA AAcg}g}|D]9}|ds|d|f0t|}|tvr|d|f`|ddkr|r|\}}}||krH|d|f|||d|f|}d}n0|d|f|||}|||d|f||||fg};|rH|\}}}|d|f|||}|H|S)Nr,rXrrrr) startswithrtr empty_tagspopr) r tag_stackmarkedrr start_name start_chunkparentsrs r rrsI F ""$$  MM3, ' ' '  '' :   MM3, ' ' '  8s?? %3<==??0 K%%NNC#5666NN6***NNC<000$F ENND+#6777NN6***$F %   tUm,,,   dE62 3 3 3FF "+--//;k*+++v  Mr c*eZdZdZdZddZdZdZdS) rya8 Represents a diffable token, generally a word that is displayed to the user. Opening tags are attached to this token when they are adjacent (pre_tags) and closing tags that follow the word (post_tags). Some exceptions occur when there are empty tags adjacent to a word, so there may be close tags in pre_tags, or open tags in post_tags. We also keep track of whether the word was originally followed by whitespace, even though we do not want to treat the word as equivalent to a similar word that does not have a trailing space.FNrCcvt||}||ng|_||ng|_||_|Sr)str__new__rqrrrw)clsr&rqrrrwobjs r rz token.__new__)sBkk#t$$#+#7xxR %.%:  "5 r c ndt|d|jd|jd|jd S)Nztoken(, ))r__repr__rqrrrwr s r rztoken.__repr__2s@ LL     t~~~t?W?W?WY Yr c t|Sr)rrs r r}z token.html6s4yyr NNrC)rrr__doc__rrrr}rr r ryrysZ  OYYYr ryc*eZdZdZ ddZdZdZdS) tag_tokenz Represents a token that is actually a tag. Currently this is just the tag, which takes up visible space just like a word but is only represented in a document by a tag. NrCct|td||||}||_||_||_|S)Nz: rv)ryrrxtagdata html_repr)rrrrrqrrrwrs r rztag_token.__new__?sSmmCD!2!2D!2!2%-&/0CEE!  r c hd|jd|jd|jd|jd|jd|jd S)Nz tag_token(rz , html_repr=z , post_tags=z , pre_tags=z, trailing_whitespace=r)rrrrqrrrwrs r rztag_token.__repr__JsG HHH III NNN MMM NNN  $ $ $ & &r c|jSr)rrs r r}ztag_token.htmlRs ~r r)rrrrrrr}rr r rr9sX555946    &&&r rceZdZdZdZdZdS) href_tokenzh Represents the href in an anchor tag. Unlike other words, we only show the href when it changes. Tc d|zS)Nz Link: %srrs r r}zhref_token.html\s T!!r N)rrrrrr}rr r rrUs4((O"""""r rTctj|r|}nt|d}t|d|}t |S)ak Parse the given HTML and returns token objects (words with attached tags). This parses only the content of a page; anything in the head is ignored, and the and elements are themselves optional. The content is then parsed by lxml, which ensures the validity of the resulting parsed document (though lxml may make incorrect guesses when the markup is particular bad). and tags are also eliminated from the document, as that gets confusing. If include_hrefs is true, then the href attribute of
tags is included as a special kind of diffable token.Tcleanup)skip_tagrQ)r iselement parse_html flatten_el fixup_chunks)r}rQbody_elrs r rRrR`sQ t1T4000 $m L L LF   r cF|rt|}t|dS)a Parses an HTML fragment, returning an lxml element. Note that the HTML will be wrapped in a
tag that was not in the original document. If cleanup is true, make sure there's no or , and get rid of any and tags. T) create_parent) cleanup_htmlr )r}rs r rrys,"D!! t4 8 8 88r z z zct|}|r||d}t|}|r|d|}t d|}|S)z This 'cleans' the HTML, meaning that any page structure is removed (only the contents of are used, if there is any and tags are removed. NrC) _search_bodyend_search_end_bodystart_replace_ins_del)r}matchs r rrsn   E "EIIKKLL! T " "E $NU[[]]N# B % %D Kr clt|}|d|||dfS)zP This function takes a word, such as 'test ' and returns ('test',' ') rN)rirstrip)wordstripped_lengths r split_trailing_whitespacers9$++--((O /! "D)9)9$: ::r c xg}d}g}|D]{}t|tr|ddkrL|d}t|d\}}td||||}g}||n=|ddkr1|d}t ||d }g}||t |r> )B5)I)I &E&UYL_```HI MM( # # # # %    U # # # #     1  ''''999xx899x"))%0000 5 /b9---..r ##I... Mr )address blockquotecenterdirdivdlfieldsetformh1h2h3h4h5h6hrisindexmenunoframesnoscriptolppretableul) dddtframesetlitbodytdtfootththeadtrc#pK|sD|jdkr(d|dt|fVnt|V|jtvr|jst |s |jsdSt|j}|D]}t|V|D]}t||Ed{V|jdkr0|dr|rd|dfV|s;t|Vt|j}|D]}t|VdSdS)a Takes an lxml element el, and generates all the text chunks for that tag. Each start tag is a chunk, each word is a chunk, and each end tag is a chunk. If skip_tag is true, then the outermost container tag is not returned (just its contents).rrkNrPrWr) rget start_tagrr&ritail split_wordsr7rend_tag)elrQr start_wordsrchild end_wordss r rrs  6U??"&&--27 7 7 7 7B--    vBGCGGBGbg&&K  $BBe=AAAAAAAAAAA v}}}M}rvvf~~&&&& $bkk((  $ $Dd## # # # # $$ $ $r z \S+(?:\s+|$)cT|r|sgSt|}|S)z_ Splits some text into words. Includes trailing whitespace on each word when appropriate. )rH _find_words)r&wordss r r2r2#s2 tzz||   E Lr z ^[ \t\n\r]cdd|jD}d|j|dS)z= The text representation of the start tag for a tag. rCc@g|]\}}d|dt|dS)rz="r0r:)r@rrs r rBzstart_tag..2sG D% *D))K&&)))r r,r.)rGattribitemsr)r4 attributess r r0r0.sY9??,,J %rv $z $ $ $$r cT|j}|rt|rdnd}d|jd|S)zg The text representation of an end tag for a tag. Includes trailing whitespace when appropriate. rrC>$  r cX|do|d S)Nr,rArErFs r rrEs( >>#   ;s~~d';';#;;r cht|d}t|t|d}|S)z Given an html string, move any or tags inside of any block-level elements, e.g. transform

word

to

word

FrT) skip_outer)r_fixup_ins_del_tagsserialize_html_fragment)r}rAs r rrHs; T5 ) ) )C "34 8 8 8D Kr ct|tr Jd|tj|dd}|rQ||ddzd}|d|d}|S|S) z Serialize a single lxml element as HTML. The serialized form includes the elements tail. If skip_outer is true, then don't serialize the outermost tag z1You should pass in an element, not a string like r}unicode)methodencodingr.rNr,)rrrtostringfindrfindrH)r4rJr}s r rLrLQs "c""DDBBBBDD " >"Vi @ @ @DDIIcNN1$%%&$TZZ__$%zz|| r ct|ddD]<}t|st||j|=dS)z?fixup_ins_del_tags that works on an lxml document in-place rr)rN)listiter_contains_block_level_tag_move_el_inside_blockrdrop_tag)rAr4s r rKrKdsj388E5))**(,,  bbf---- r c.|jtD]}dSdS)zPTrue if the element contains any block-level elements, like

, , etc. TF)rVany_block_level_tag)r4s r rWrWps)bg*+tt 5r c|j}|jtD]}||urnK ||}|j|_d|_|t||g|dd<dSt |D]}t |rKt|||jr3||}|j|_d|_| |\||}| ||| ||jr6||}|j|_d|_| d|dSdS)zt helper for _fixup_ins_del_tags; actually takes the etc tags and moves them inside any block-level tags. Nr) makeelementrVr[r&rrUrWrXr1addnextr4rtr) r4rr]block_level_el children_tagr6tail_tag child_tagtext_tags r rXrXys.K!"'#67    # # E $#{3'' G DHH%%%111b $ $ $U + + $ !% - - -z (&;s++ %  !  h'''# C((I JJui ( ( (   U # # # # w;s##  !X r c|}|j}|j}|r4t|s|pd|z}n|djpd|z|d_||}|r9|}||jpd|z|_n|jpd|z|_||||dz<dS)z Removes an element, but merges its contents into its place, e.g., given

Hi there!

, if you remove the element you get

Hi there!

rCrpNr) getparentr&r1rir getprevious getchildren)r4parentr&r1rpreviouss r _merge_element_contentsrjs \\^^F 7D 7D 52ww 5JB$&DDb6;,"4BrFK LL  E 9>>##  !;,"4FKK%]0bD8HMNN,,F5q=r c<eZdZdZdZejdefdZdS)rZzt Acts like SequenceMatcher, but tries not to find very small equal blocks amidst large spans of changes r-r(ctt|jt|j}|jt|dzt j|}fd|DS)Nr1c<g|]}|dks|d|S)r-r)r@r thresholds r rBzBInsensitiveSequenceMatcher.get_matching_blocks..s=   7Y&&Aw'&&&r )minrirXrnrget_matching_blocks)r sizeactualrns @r rpz.InsensitiveSequenceMatcher.get_matching_blockssv"%c$&kk3tv;;"?"?'+~  419--  4T::        r N) rrrrrncythonr rUrprr r rZrZsL I \ T   \   r rZ__main__) _diffcommand)r%)F)T)brs ImportErrorrrCrdifflibinspect isfunctionget_close_matches"cython.cimports.lxml.html._difflibrr itertools functoolsoperatorrelxmlr lxml.htmlr r!__all__partialgroupby itemgetterrr rrr7r}r8r<r"r?rDr\rErUrsrFr#rrrrrrrrrryrrrRrcompileISsearchrrsubrrrrr frozensetrblock_level_tagsblock_level_container_tagssortedr[rUfindallr9r2rrBr0r3rr rrrLrKrWrXrjfinalrrZrrumainrr r rs MMMM::::::::[]]FFF ,%%%%%%NNNw'344\k Z\\ \BBBBBBB,,,NNN ?,++++++,  )))))) J ''i' (9?Rx?RST?U?UVVVcU_b&+******$$$#1&#&#&#&#P - - -111         # # #"&&&8%%%P####$<SS$4444444                $H4HHHHV2t2222jCB8""""""""    2 9 9 9 9rz,RT 229 2:mRT"$Y77>2:1249==A   ;;;222lV^It 7 7 !6>)YY888..6,V^Iyy B B B 8 8  %fnUEE&&113333-4-4 $$$$6bj"$//7 # =117%%%!!!###   <<<&   F---0        & z&&&&&&Ls&&0AA43A4