str.replace, es sei denn, auf string folgt ein bestimmter Text

Oct 25 2020

Überarbeitung der vorherigen Frage:

Wie kann ich alle "," (dh Komma dann Leerzeichen) durch "_" ersetzen, außer wenn auf "," (Komma dann Leerzeichen) das Wort "LLC" oder "Inc" folgt (dann nichts tun)?

Ich will es verändern:

  1. "TEXAS ENERGY MUTUAL, LLC, BOBBY GILLIAM, STEVE PEREIRA und ANDY STITT"
  2. "Grape, LLC, Andrea Gray, Jack Smith"
  3. "Stephen Winters, Apple, Birne, Inc, Sarah Smith"

Dazu:

  1. "TEXAS ENERGY MUTUAL, LLC_BOBBY GILLIAM_STEVE PEREIRA_ANDY STITT"
  2. "Grape, LLC_Andrea Gray_Jack Smith"
  3. "Stephen Winters_Apple_pear, Inc_Sarah Smith"

Ich dachte, es würde mit einer Variation des folgenden Codes beginnen, aber ich kann die Ausnahmebedingungen nicht herausfinden.

df ['Column_Name'] = df ['Column_Name']. str.replace (',', '_') Prost!

Antworten

1 ernest_k Oct 25 2020 at 05:35

Sie können einen regulären Ausdruck durch einen negativen Lookahead ersetzen :

#no idea why Inc|LLC or LLC|Inc will skip the first
df['Column_Name'].str.replace(', (?!=|Inc|LLC)', '_')

Ausgabe:

0    TEXAS ENERGY MUTUAL, LLC_BOBBY GILLIAM_STEVE P...
1                    Grape, LLC_Andrea Gray_Jack Smith
2          Stephen Winters_Apple_pear, Inc_Sarah Smith
Name: ColumnName, dtype: object

1 chai Oct 25 2020 at 05:34

Verwenden Python Regex Modul re für mit dem Muster , (?!Inc|LLC)alle Vorkommen zu finden , , ohne auf IncoderLLC

import re

strings = ["Banana, orange", "Grape, LLC", "Apple, pear, Inc"]

[re.sub(", (?!Inc|LLC)",'_',string) for string in strings]
#['Banana_orange', 'Grape, LLC', 'Apple_pear, Inc']
DozParp Oct 25 2020 at 05:52

der einfache Weg:

def replace(str):
   x = str.split(', ')
   buf = x[0]
   for i in range(1, len(x)): 
      if x[i].startswith('LLC'):
         buf += ', ' + x[i]
      elif x[i].startswith('Inc'):
         buf += ', ' + x[i]
      else:
         buf += '_' + x[i]
   return buf

und dann versuchen replace('a, b, LLC, d')