9,jMTdZddlmZdZdgZddlmZddlZddlm Z m Z m Z m Z m Z mZmZmZmZmZmZddlmZmZmZmZmZmZdd lmZmZdd lmZm Z m!Z!m"Z"dd l#m$Z$erdd l%m&Z&dd lm'Z'ddl(m)Z)m*Z*m+Z+dZ,e e e-e-fe-e-gdfZ.GddeeZ/Gdde!Z0dS)zCUse the HTMLParser library to parse HTML files that aren't too bad.) annotationsMITHTMLParserTreeBuilder) HTMLParserN) AnyCallablecastDictIterableListOptional TYPE_CHECKINGTupleTypeUnion) AttributeDictCDataComment DeclarationDoctypeProcessingInstruction)EntitySubstitution UnicodeDammit)DetectsXMLParsedAsHTMLHTMLHTMLTreeBuilderSTRICTParserRejectedMarkup) BeautifulSoup)NavigableString) _Encoding _Encodings _RawMarkupz html.parserceZdZUdZded<dZded< edd.dZd ed <ded<ded<d/dZd0dZ d1d2dZ d1d3dZ d4d Z e j d!Ze j d"Zed5d%Zd6d&Zd6d'Zd4d(Zd7d*Zd4d+Zd4d,Zd-S)8BeautifulSoupHTMLParserreplacestrREPLACEignoreIGNOREon_duplicate_attributesoupr argsrr-&Union[str, _DuplicateAttributeHandler]kwargsc||_||_|jj|_t j|g|Ri|g|_|dSN)r.r-builderattribute_dict_classr__init__already_closed_empty_element_initialize_xml_detector)selfr.r-r/r1s GC:\PYTHON\GemmaClient\venv\Lib\site-packages\bs4/builder/_htmlparser.pyr6z BeautifulSoupHTMLParser.__init__Usc &<#$(L$E!D242226222-/) %%'''''z List[str]r7messagereturnNonec t|r3r)r9r<s r:errorzBeautifulSoupHTMLParser.errorps#7+++r;tagattrsList[Tuple[str, Optional[str]]]cd|||d||ddS)zHandle an incoming empty-element tag. html.parser only calls this method when the markup looks like . F)handle_empty_elementcheck_already_closedN)handle_starttag handle_endtag)r9rArBs r:handle_startendtagz*BeautifulSoupHTMLParser.handle_startendtags@ S%eDDD 3U;;;;;r;TrEboolc4|}|D]Y\}}|d}||vrG|j}||jkr |d|jfvr|||<1t t |}||||T|||<Z|jjjr| \}} ndx}} |j |dd||| } | :| j r3|r1| |d|j ||j||dSdS)zHandle an opening tag, e.g. '' :param handle_empty_element: True if this tag is known to be an empty-element tag (i.e. there is not expected to be any closing tag). N) sourceline sourceposFrF)r5r-r+r)r _DuplicateAttributeHandlerr.r4store_line_numbersgetposrHis_empty_elementrIr7append_root_tag_name_root_tag_encountered) r9rArBrE attr_dictkeyvalueon_duperNrOtagObjs r:rHz'BeautifulSoupHTMLParser.handle_starttagsv$(#<#<#>#>  ' 'JC}i5dk))t| 444%*IcNN"#=wGGGGIsE2222!& # 9  / *$(KKMM !J %) )J** tY:+    &"9 >R    s  ? ? ?  - 4 4S 9 9 9   &  & &s + + + + + ' &r;rGc|r%||jvr|j|dS|j|dS)zHandle a closing tag, e.g. '' :param tag: A tag name. :param check_already_closed: True if this tag is expected to be the closing portion of an empty-element tag, e.g. ''. N)r7remover.rI)r9rArGs r:rIz%BeautifulSoupHTMLParser.handle_endtagsT  )C4+L$L$L  - 4 4S 9 9 9 9 9 I # #C ( ( ( ( (r;datac:|j|dS)z4Handle some textual data that shows up between tags.N)r. handle_datar9r^s r:r`z#BeautifulSoupHTMLParser.handle_datas d#####r;z ^([0-9]+)(.*)z^([0-9a-f]+)(.*)nameTuple[str, bool, str]cd}d}d}d}|j}|ds|dr|dd}d}|j}d} t||}ni#t$r\||}|Bt|d |}|d}YnwxYw|d}|}ntj|\}}|||fS) aConvert a numeric character reference into an actual character. :param name: The number of the character reference, as obtained by html.parser :return: A 3-tuple (dereferenced, replacement_added, extra_data). `dereferenced` is the dereferenced character reference, or the empty string if there was no reference. `replacement_added` is True if the reference could only be dereferenced by replacing content with U+FFFD REPLACEMENT CHARACTER. `extra_data` is a portion of data following the character reference, which was deemed to be normal data and not part of the reference at all. rMF xXNr) &_DECIMAL_REFERENCE_WITH_FOLLOWING_DATA startswith"_HEX_REFERENCE_WITH_FOLLOWING_DATAint ValueErrorsearchgroupsrnumeric_character_reference) clsrb dereferencedreplacement_added extra_databasereg real_namematchs r:(_dereference_numeric_character_referencez@BeautifulSoupHTMLParser._dereference_numeric_character_references  !& 8 ??3   94??3#7#7 98DD8C"&  /D$II / / /JJt$$E  q 1488 "\\^^A.  /"  LJJ.;.WXa.b.b +L+. ::sA!!A#CCc||\}}}|r d|j_||||||dSdS)zHandle a numeric character reference by converting it to the corresponding Unicode character and treating it as textual data. :param name: Character number, possibly in hexadecimal. TN)rzr.contains_replacement_charactersr`)r9rbrsrtrus r:handle_charrefz&BeautifulSoupHTMLParser.handle_charref"st7;6c6cdh6i6i3 '  =8)rAr(rBrCr=r>)T)rAr(rBrCrErKr=r>)rAr(rGrKr=r>)r^r(r=r>)rbr(r=rc)rbr(r=r>)rr(r=r>)__name__ __module__ __qualname__r)__annotations__r+r6r@rJrHrIr`recompilerjrl classmethodrzr}rrrrrr;r:r&r&>sGF $JQ ((((((.CBBB++++,,,, <<<<0&* <,<,<,<,<,|)))))$$$$$.8RZ-H-H*)34F)G)G&4;4;4;[4;l ) ) ) )&########    111111r;r&ceZdZUdZdZded<dZded<eZded<ee e gZ d ed <d ed <dZ ded < d#d$fd Z d%d&dZefd'd"ZxZS)(rzA Beautiful soup `bs4.builder.TreeBuilder` that uses the :py:class:`html.parser.HTMLParser` parser, found in the Python standard library. FrKis_xmlT picklabler(NAMEz Iterable[str]featuresz$Tuple[Iterable[Any], Dict[str, Any]] parser_argsTRACKS_LINE_NUMBERSNOptional[Iterable[Any]] parser_kwargsOptional[Dict[str, Any]]r1rc t}dD] }||vr||}|||<!tt|jdi||pg}|pi}||d|d<||f|_dS)aConstructor. :param parser_args: Positional arguments to pass into the BeautifulSoupHTMLParser constructor, once it's invoked. :param parser_kwargs: Keyword arguments to pass into the BeautifulSoupHTMLParser constructor, once it's invoked. :param kwargs: Keyword arguments for the superclass constructor. r,Fconvert_charrefsNr)dictpopsuperrr6updater)r9rrr1extra_parser_kwargsargrY __class__s r:r6zHTMLParserTreeBuilder.__init__s$#ff. 1 1Cf}} 3+0#C(3#T**3==f===!'R %+ 0111,1 ()'7r;markupr$user_specified_encodingOptional[_Encoding]document_declared_encodingexclude_encodingsOptional[_Encodings]r=DIterable[Tuple[str, Optional[_Encoding], Optional[_Encoding], bool]]c#8Kt|tr |dddfVdSg}|r||g}|r||t|||d|}|jt d|j|j|j|jfVdS)a2Run any preliminary steps necessary to make incoming markup acceptable to the parser. :param markup: Some markup -- probably a bytestring. :param user_specified_encoding: The user asked to try this encoding. :param document_declared_encoding: The markup itself claims to be in this encoding. :param exclude_encodings: The user asked _not_ to try any of these encodings. :yield: A series of 4-tuples: (markup, encoding, declared encoding, has undergone character replacement) Each 4-tuple represents a strategy for parsing the document. This TreeBuilder uses Unicode, Dammit to convert the markup into Unicode, so the ``markup`` element of the tuple will always be a string. NFT)known_definite_encodingsuser_encodingsis_htmlrzPCould not convert input to Unicode, and html.parser will not accept bytestrings.) isinstancer(rTrunicode_markuproriginal_encodingdeclared_html_encodingr|)r9rrrrrrdammits r:prepare_markupz$HTMLParserTreeBuilder.prepare_markups2 fc " " 4u- - - - F57 " E % + +,C D D D*, % >  ! !"< = = = %=)/      ('b  %(-6      r; _parser_classtype[BeautifulSoupHTMLParser]r>c"|j\}}t|tsJ|jJ||jg|Ri|} |||n!#t $r}t|d}~wwxYwg|_dS)z :param markup: The markup to feed into the parser. :param _parser_class: An HTMLParser subclass to use. This is only intended for use in unit tests. N) rrr(r.feedcloseAssertionErrorrr7)r9rrr/r1parseres r:rzHTMLParserTreeBuilder.feeds ' f&#&&&&& y$$$ty:4:::6:: * KK    LLNNNN * * *'q)) )  * /1+++s)A'' B1BB)NN)rrrrr1r)NNN) rr$rrrrrrr=r)rr$rrr=r>)rrr__doc__rrr HTMLPARSERrrrrrr6rr&r __classcell__)rs@r:rrqs FID#T62H22225555!%$$$$04268888888B8<:>26 FFFFFPUl111111111r;)1r __future__r __license____all__ html.parserrrtypingrrr r r r r rrrr bs4.elementrrrrrr bs4.dammitrr bs4.builderrrrrbs4.exceptionsrbs4r r! bs4._typingr"r#r$rr(rPr&rrr;r:rsII""""""  #"""""                           988888880/////!!!!!!++++++  %tCH~sC&@$&FGp1p1p1p1p1j*@p1p1p1f T1T1T1T1T1OT1T1T1T1T1r;